Methods, apparatuses, and devices for training an information extraction model and obtaining a knowledge graph

By using the meta-learning method to update the information extraction model multiple times, the fitting of noise data is reduced, and the problem of poor robustness of the information extraction model to noise data in the existing technology is solved, and more efficient and accurate knowledge graph construction is achieved.

CN111737552BActive Publication Date: 2025-06-10INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010500623.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-04
Publication Date
2025-06-10
Estimated Expiration
2040-06-04

AI Technical Summary

Technical Problem

The existing information extraction model is poorly robust to noise data, which is easy to fit noise data, resulting in performance losses.

Method used

By obtaining a training sample set containing non-noise samples and noise samples, the initial information extraction model is updated multiple times, the meta-learning method is used to simulate the impact of error labels, and the model parameters are gradually adjusted to reduce the fitting of noise data, and a target information extraction model that does not fit the noise data is obtained.

Benefits of technology

It improves the robustness of the information extraction model, ensures that the model performs better in noise, and the built knowledge graph is more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111737552B_ABST
    Figure CN111737552B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technologies, and provides a method, an apparatus, and a device for training an information extraction model and obtaining a knowledge graph, so as to improve the robustness of the information extraction model. The method includes: training an initial first information extraction model, updating the first information extraction model based on the prediction result of a noise sample to obtain a first intermediate model; updating the first information extraction model based on the difference between the prediction result of the first intermediate model for the noise sample and the prediction result of an initial second information extraction model for a non-noise sample to obtain a second intermediate model; updating the second intermediate model based on the prediction result of the second intermediate model for the non-noise sample to obtain a reference model; adjusting the parameters of the reference model based on a preset smoothing coefficient to obtain a target information extraction model. The present application updates the model parameters based on a meta-learning method, and the updated model is more robust and the constructed knowledge graph is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Today, with the rapid development of the Internet, it is an era dominated by information and big data. It is very important to extract the content of interest in the era of information explosion. Information extraction research aims to provide people with more powerful information acquisition tools to cope with the severe challenges brought by information explosion. Currently, the most commonly used method is remote supervision relation extraction, which maps the existing knowledge base to rich unstructured data to generate a large amount of training data, and then trains the information extraction model. However, remote supervision inevitably has the problem of data annotation errors, generating noisy data with incorrect labels. Since neural networks are easily overfitted to noisy data, training the information extraction model based on this noisy data will seriously affect the performance of the model.

[0003] In summary, the current information extraction model has poor robustness to noisy data, is easily overfitted to noisy data, resulting in performance loss of the information extraction model. Summary of the Invention

[0004] Embodiments of the present application provide a method, device, and equipment for training an information extraction model and obtaining a knowledge graph, so as to improve the robustness of the information extraction model and further improve the accuracy of the knowledge graph.

[0005] A method for training an information extraction model provided by an embodiment of the present application includes:

[0006] Obtain a training sample set including a non-noise sample set and a noise sample set, where each non-noise sample in the non-noise sample set is a sentence packet with a labeled non-noise label, and each noise sample in the noise sample set is a sentence packet with a labeled noise label. Each sentence packet includes multiple sentences for describing the association information between entities, the sentences are unstructured natural language texts, and the association information includes the relationship between entities or events involved by entities;

[0007] Train an initial first information extraction model using the noise sample set, and during the training process, update the first information extraction model based on the difference between the prediction result of the noise sample by the first information extraction model and the noise label corresponding to the noise sample, to obtain a first intermediate model;

[0008] Update the first information extraction model based on the difference between the prediction result of the noise sample by the first intermediate model and the prediction result of the non-noise sample by the initial second information extraction model, to obtain a second intermediate model, where the first information extraction model and the second information extraction model have the same parameters, and the second information extraction model is used to guide the training of the first information extraction model;

[0009] Update the second intermediate state model based on the difference between the prediction result of the non-noise samples by the second intermediate state model and the non-noise labels corresponding to the non-noise samples, to obtain a trained reference model;

[0010] Adjust the parameters of the trained reference model based on a preset smoothing coefficient, to obtain a target information extraction model for obtaining entity association information, where the entity association information is used to construct a knowledge graph.

[0011] An apparatus for training an information extraction model provided by an embodiment of the present application includes:

[0012] An acquisition unit, configured to acquire a training sample set including a non-noise sample set and a noise sample set, where each non-noise sample in the non-noise sample set is a sentence pack with a labeled non-noise label, each noise sample in the noise sample set is a sentence pack with a labeled noise label, each sentence pack includes multiple sentences for describing the association information between entities, the sentences are unstructured natural language texts, and the association information includes the relationship between entities or the events involved by the entities;

[0013] A first update unit, configured to train an initial first information extraction model using the noise sample set, and update the first information extraction model based on the difference between the prediction result of the noise samples by the first information extraction model and the noise labels corresponding to the noise samples during the training process, to obtain a first intermediate state model;

[0014] A second update unit, configured to update the first information extraction model based on the difference between the prediction result of the noise samples by the first intermediate state model and the prediction result of the non-noise samples by an initial second information extraction model, to obtain a second intermediate state model, where the first information extraction model and the second information extraction model have the same parameters, and the second information extraction model is used to guide the training of the first information extraction model;

[0015] A third update unit, configured to update the second intermediate state model based on the difference between the prediction result of the non-noise samples by the second intermediate state model and the non-noise labels corresponding to the non-noise samples, to obtain a trained reference model;

[0016] An adjustment unit, configured to adjust the parameters of the trained reference model based on a preset smoothing coefficient, to obtain a target information extraction model for obtaining entity association information, where the entity association information is used to construct a knowledge graph.

[0017] Optionally, the noise sample set includes M, and the acquisition unit is specifically configured to:

[0018] For each non-noise sample in the non-noise sample set, perform M label transfers on each non-noise sample respectively according to the non-noise labels of other non-noise samples except itself, to obtain M noise labels corresponding to each non-noise sample, where M is a positive integer;

[0019] Generate M noise samples corresponding to each non-noise sample respectively based on the sentence packets included in each non-noise sample and the M noise labels corresponding to each non-noise sample. The sentence packets of the noise samples corresponding to each non-noise sample are the same, and the noise samples obtained by each non-noise sample for each label transfer belong to a noise sample set.

[0020] Optionally, the obtaining unit is specifically configured to:

[0021] For any non-noise sample, perform the following process each time a noise label is obtained by label transfer:

[0022] Obtain the similarity between the non-noise sample and other non-noise samples except the non-noise sample, and use the non-noise label of any one of the top N other non-noise samples with the similarity sorted from high to low as the noise label obtained after label transfer of the non-noise sample, where N is a positive integer.

[0023] Optionally, the prediction result of the first information extraction model for the noise sample is the first prediction label; the first updating unit is specifically configured to:

[0024] Input the noise samples in the M noise sample sets into the first information extraction model in batches, and obtain M batches of first prediction labels output by the first information extraction model, where the noise samples in the same batch belong to the same noise sample set, and the noise samples in different batches belong to different noise sample sets;

[0025] Obtain a first classification loss function determined based on the difference between each batch of first prediction labels and the noise labels of the corresponding noise samples;

[0026] Perform a gradient update on the first information extraction model once according to each first classification loss function, to obtain M first intermediate models;

[0027] The second updating unit is specifically configured to: obtain a fusion difference based on the difference between the prediction result of each first intermediate model for the noise sample and the prediction result of the second information extraction model for the non-noise sample, and update the first information extraction model according to the fusion difference to obtain the second intermediate model.

[0028] Optionally, the prediction result of the first intermediate state model for the noise samples is a second prediction label, and the prediction result of the second information extraction model for the non-noise samples is a third prediction label; specifically, the second updating unit is configured to:

[0029] Input the noise samples in the M noise sample sets into the corresponding first intermediate state models respectively, to obtain the second prediction labels output by each first intermediate state model; and input the non-noise samples in the non-noise sample set into the second information extraction model, to obtain the third prediction labels output by the second information extraction model;

[0030] Determine a consistency loss function based on KL divergence respectively according to the differences between each batch of the second prediction labels and the third prediction labels;

[0031] Average the M consistency loss functions to obtain the fusion difference;

[0032] Perform at least one gradient update on the first information extraction model according to the fusion difference, to obtain the second intermediate state model, where the error between the prediction result of the second intermediate state model and the prediction result of the second information extraction model is within a specified range.

[0033] Optionally, the prediction result of the second intermediate state model for the non-noise samples is a fourth prediction label; specifically, the third updating unit is configured to:

[0034] Input the non-noise samples in the non-noise sample set into the second intermediate state model, to obtain the fourth prediction labels output by the second intermediate state model;

[0035] Obtain a second classification loss function determined according to the difference between the fourth prediction label and the non-noise label of the corresponding non-noise sample;

[0036] Perform at least one gradient update on the second intermediate state model according to the second classification loss function, to obtain the trained reference model, where the error between the prediction result of the trained reference model and the prediction result of the second intermediate state model is within a specified range.

[0037] Optionally, the adjusting unit is specifically configured to:

[0038] Perform exponential averaging on the parameters of the trained reference model according to the preset smoothing coefficient, to obtain the target information extraction model.

[0039] A method for obtaining a knowledge graph provided by an embodiment of the present application includes:

[0040] Obtain a text to be processed, where the text to be processed is an unstructured natural language text for describing the association information between entities;

[0041] Input the text to be processed into the trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model, where the target information extraction model is obtained by training using any of the above methods for training an information extraction model, and the association information includes the relationship between entities or the events involved by entities.

[0042] Construct a knowledge graph based on the entity association information.

[0043] An apparatus for obtaining a knowledge graph provided by an embodiment of the present application includes:

[0044] An acquisition unit, configured to acquire the text to be processed, where the text to be processed is an unstructured natural language text for describing the association information between entities, and the association information includes the relationship between entities or the events involved by entities;

[0045] An information extraction unit, configured to input the text to be processed into the trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model, where the target information extraction model is obtained by training using any of the above methods for training a target information extraction model;

[0046] A construction unit, configured to construct a knowledge graph based on the entity association information.

[0047] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the above method for training an information extraction model or the steps of the method for obtaining a knowledge graph.

[0048] A computer-readable storage medium provided by an embodiment of the present application includes program code, and when the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps of the above method for training an information extraction model or the steps of the method for obtaining a knowledge graph.

[0049] The embodiments of the present application provide a method, apparatus, electronic device, and storage medium for training an information extraction model and obtaining a knowledge graph. When training the information extraction model, the training sample data set used contains both non-noise samples and noise samples, and these samples are all unstructured natural language texts. Therefore, the target information extraction model trained based on these samples can accurately identify entities and entity relationships, or events involved by entities in the text to be processed, and can then be used to construct a more accurate knowledge graph. During the model training process, first, the first information extraction model is trained using the noise sample set to simulate the training process of the first information extraction model under noise data. Then, based on the difference between the prediction results of the updated first intermediate model for the noise samples and the prediction results of the initial second information extraction model for the non-noise samples, the initial first information extraction model is updated to ensure that the updated second intermediate model can give the same prediction results as the initial second information extraction model, thereby ensuring that the updated second intermediate model does not overfit the noise data. Finally, the second intermediate model is updated by the method of stochastic gradient descent to obtain the trained reference model. Since the trained reference model obtained here does not overfit the noise data, the target information extraction model obtained by adjusting the parameters of the reference model also does not overfit the noise data, and the model performance is better, and it performs more robustly in the presence of noise. Based on this model, the text to be processed can be processed to obtain more accurate entity relationships, and thus the constructed knowledge graph is also more accurate.

[0050] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0052] Figure 1 It is a schematic structural diagram of the first information extraction model in the embodiments of the present application;

[0053] Figure 2 It is an optional schematic diagram of an application scenario in the embodiments of the present application;

[0054] Figure 3 It is an optional flowchart of a method for training an information extraction model in the embodiments of the present application;

[0055] Figure 4 A robustness training framework for a remote supervision relation extraction model based on meta - learning in the embodiments of the present application;

[0056] Figure 5A A structural schematic diagram of the second relation extraction model in the embodiments of the present application;

[0057] Figure 5B A structural schematic diagram of the third relation extraction model in the embodiments of the present application;

[0058] Figure 6 A schematic diagram of an experimental result in the embodiments of the present application;

[0059] Figure 7A An optional method flowchart for obtaining a knowledge graph in the embodiments of the present application;

[0060] Figure 7B A schematic diagram of a framework for using a relation extraction model in the embodiments of the present application;

[0061] Figure 8 A complete method implementation time - sequence flowchart for an optional training information extraction model in the embodiments of the present application;

[0062] Figure 9 A composition structural schematic diagram of a device for training an information extraction model in the embodiments of the present application;

[0063] Figure 10 A composition structural schematic diagram of a device for obtaining a knowledge graph in the embodiments of the present application;

[0064] Figure 11 A composition structural schematic diagram of an electronic device in the embodiments of the present application;

[0065] Figure 12 A hardware composition structural schematic diagram of a computing device applying the embodiments of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0067] Some concepts involved in the embodiments of the present application are introduced below.

[0068] Unstructured text: Unstructured data in the form of text (such as characters, numbers, punctuation marks, various printable symbols, etc.); typical representatives of unstructured or semi-structured text data are documents in a library database, which may contain structured fields such as title, author, publication date, length, classification, etc., and may also contain a large number of unstructured text components such as abstracts and body content.

[0069] Natural language: Usually refers to a language that naturally evolves with culture. For example, English, Chinese, and Japanese are examples of natural languages, while Esperanto is an artificial language, that is, a language created for certain specific purposes. However, sometimes all languages used by humans (including the above-mentioned languages that naturally evolve with culture, as well as artificial languages) are regarded as natural languages, as opposed to artificial languages designed for computers such as programming languages.

[0070] Distant supervision: It is a method of data augmentation for data annotation based on an existing knowledge base, mainly to solve the problem of insufficient data, but there may also be problems with incorrect data annotation. Distant supervision is a semi-supervised learning algorithm. In the embodiments of the present application, the non-noise sample set in the training sample set can be used to perform data annotation on the knowledge base through distant supervision to obtain multiple sentence packets with labeled non-noise tags.

[0071] Knowledge Graph: Describes concepts, entities, and their relationships in the objective world in a structured form, expresses the information on the Internet in a form closer to the human cognitive world, and provides an ability to better organize, manage, and understand the vast amount of information on the Internet. The Knowledge Graph is to discover the connections between all things in the world. Technically, the data is stored in the form of triples, each containing a subject, a relation, and an object. In the embodiments of the present application, based on the trained target information extraction model, entities and entity association information in natural language text can be obtained. Entity association information includes the relationships between entities or the event information involved by entities. The relationship between entities can be represented by triples, and then a graph database (such as neo4j, etc.) can be used for knowledge storage to construct a Knowledge Graph. Examples of triple information are: (A, B, teacher-student relationship). The event information involved by entities and entities can also be represented by similar triples or more-tuple information and further construct a Knowledge Graph. Examples of triple information are: (C, a certain year, sports meeting gold medal), examples of quadruple information are: (C, a certain year, a certain city, sports meeting gold medal), and examples of quintuple information are: (C, a certain year, a certain city, sports meeting, diving, gold medal), etc.

[0072] Relation extraction: Automatically extracting entities and identifying the associated information of entities from unstructured text to be processed. It belongs to the category of natural language understanding and is also a method for constructing and expanding knowledge graphs. It is an important research topic in the field of information extraction. Its main purpose is to extract the associated information of marked entities in a sentence, such as the semantic relationship between marked entity pairs. That is, on the basis of entity recognition, determine the relationship category between entity pairs in unstructured text and form structured data for storage and retrieval. Information extraction is a very important part of natural language processing, including entity extraction, relation extraction between entities, and event extraction involving entities. The model training method in the embodiments of this application can be applied to relation extraction tasks and can also be applied to event extraction tasks.

[0073] Sentence pack: It refers to a pack composed of multiple sentences, and each sentence is used to describe the relationship between entities or events involving entities. Here, an entity refers to an objectively existing and distinguishable thing, which can be a specific person, thing, or object, or an abstract concept or relationship. For example, for the sentence: The student of Zhang is Li, and his height is 180 cm. The entities in this sentence refer to people, namely Zhang and Li, and the relationship between the entities is the teacher-student relationship. For the sentence: A concert was held in place A. The entity in this sentence refers to the location: A, and the event involving the entity is the concert.

[0074] Sample labels, non-noise sample labels, and noise sample labels: Sample labels are used to represent entities that are samples and the entity-related information. The entity-related information includes the type of relationship between entities, or the type of event that the entity is involved in, etc. In the relationship extraction task, the label represents the relationship between entities in the sample, such as father-son, mother-son, teacher-student, etc. In addition, each label also corresponds to a corresponding probability value, which represents the probability that the extracted entity relationship belongs to the relationship type represented by the label. For example, the probability that the relationship between a pair of entities belongs to the father-son relationship; in the event extraction task, the label represents the type of event that the entity is involved in, such as concert, sports meeting, etc. Similarly, the label also corresponds to a corresponding probability value, which represents the probability that the event involved by the extracted entity belongs to the event type represented by the label. For example, the probability of belonging to a concert. In the embodiments of the present application, many labels are listed. For training samples, they include noise labels and non-noise labels. Among them, non-noise labels are those obtained by data relabeling through remote supervision, while noise labels are wrong labels directly obtained by label transfer of non-noise samples. The non-noise label annotates the entity, entity-related information, and the probability corresponding to the entity-related information of the non-noise sample, while the noise label includes the entity, entity-related information, and the probability corresponding to the entity-related information of the noise sample. In the embodiments of the present application, both noise labels and non-noise labels are used as training data, and the corresponding probability is usually 1. That is, during training, both noise samples and non-noise samples are regarded as positive samples, and they have different roles in the process of training different models. The noise labels annotated by noise samples are incorrect and are used to simulate the training process of the model under incorrect noise data. Taking the noise sample as an example, the corresponding noise label can be expressed as (A, B, father-son relationship, 1), indicating that entity A and entity B are in a father-son relationship. The father-son relationship is the relationship between entities, that is, the entity-related information, and the corresponding probability value is 1; taking the non-noise sample as an example, the corresponding non-noise label can be expressed as (A, B, mother-son relationship, 1), indicating that entity A and entity B are in a mother-son relationship. The mother-son relationship is used as the entity-related information, and the corresponding probability value is 1.

[0075] Predicted Labels: During the model training process of this application, the model will obtain multiple labels predicted at different training stages based on samples, which are called predicted labels. The multiple predicted labels are respectively represented as: the first predicted label, the second predicted label, the third predicted label, and the fourth predicted label. These predicted labels are the predicted labels of the samples obtained through model prediction. The model in the embodiments of this application is similar to a multi-classifier. Through the model, multiple probability values can be predicted. Different probability values correspond to different entity association information, and the sum of the probability values is 1. The model will finally take the maximum probability value and the corresponding entity association information as the prediction result. The predicted labels output by the model also include entities, entity association information, and the probabilities corresponding to the entity association information. For example, (A, B, parent-child relationship, 0.95). In addition, the predicted labels can also include various entity association information and the corresponding probabilities. The probability values obtained by prediction here can range from 0 to 1. In the embodiments of this application, by comparing these predicted labels with the sample labels, the differences between the entities, entity association information, and the probabilities corresponding to the entity association information included in the predicted labels and those included in the sample labels are calculated, and the model parameters are adjusted according to the differences.

[0076] Meta Learning: A type in the field of machine learning algorithms, mainly to enable machines to learn to learn, which means using past knowledge and experience to guide the learning of new tasks, having the ability to learn to learn, and has been widely applied to tasks such as few-shot learning and domain transfer. Meta learning can effectively solve the problems of few-shot learning, long-tail distribution, and domain adaptation in the NLP (Natural Language Processing) field. In the embodiments of this application, the parameters of the model are updated based on the meta learning method, making the updated model more robust.

[0077] KL (Kullback-Leibler) Divergence: Also known as relative entropy or information divergence, it is an asymmetric measure of the difference between two probability distributions. In information theory, relative entropy is equivalent to the difference in the information entropy of two probability distributions. Relative entropy is the loss function of some optimization algorithms, such as the Expectation-Maximization algorithm (EM). At this time, one of the probability distributions participating in the calculation is the true distribution, and the other is the theoretical (fitted) distribution. Relative entropy represents the information loss generated when using the theoretical distribution to fit the true distribution.

[0078] Information extraction model: The information extraction model in the embodiments of the present application can be a relation extraction model or an event extraction model. The relation extraction model is mainly used to extract entity and entity relationships in the text, while the event extraction model is mainly used to extract entities and events involved by entities in the text. Among them, the first information extraction model, the second information extraction model, the first intermediate state model, the second intermediate state model, and the reference model listed in the embodiments of the present application are information extraction models obtained at different training stages, and the difference is that the parameters of these models are different. The initial first information extraction model is randomly initialized before model training, and the parameters of the initial second information extraction model are the same as those of the initial first information extraction model. The second information extraction model is used to guide the training of the first information extraction model, and other models are trained based on different sample data.

[0079] Student model and teacher model: In the training of neural network models, the teacher model is used to guide the training of the student model. Generally, the prediction ability of the teacher model is much higher than that of the student model. Therefore, training the student model based on the teacher model can improve the robustness of the student model. In the embodiments of the present application, the student model includes the initial first information extraction model, the first intermediate state model, the second intermediate state model, and the reference model, and the teacher model includes the initial second information extraction model and the finally obtained target information extraction model. The parameters of the teacher model are obtained by exponential averaging of the parameters of the learning model. Exponential averaging is a self-enhancing method. Therefore, the teacher model obtained by self-enhancing the student model has better prediction ability.

[0080] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0081] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0082] Among them, Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0083] Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural languages. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural languages, that is, the languages commonly used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0084] The target information extraction model proposed in the embodiments of this application is mainly applied to process natural language texts. The training and usage methods of the target information extraction model can be divided into two parts, including a training part and an application part; among them, the training part involves the technical field of machine learning. In the training part, the target information extraction model is trained through machine learning technology, so that after the samples in the training sample set pass through the target information extraction model, the entities in the samples, as well as the relationships between entities or the events involved by the entities, are extracted, and the model parameters are continuously adjusted through optimization algorithms to obtain the trained target information extraction model; the application part is used to perform entity extraction, relationship extraction, event extraction, etc. on the natural language text to be processed by using the target information extraction model trained in the training part, and identify the entities in the natural language text, the entity relationships between entities, or the events involved by the entities, etc.

[0085] The following briefly introduces the design concept of the embodiments of this application:

[0086] Taking relation extraction in information extraction as an example, relation extraction is a very important branch in the field of natural language processing and a key technology for constructing knowledge graphs. Relation extraction refers to automatically identifying the semantic relations between entities in a sentence containing two entities. Due to the lack of labeled training data, remote supervision and deep learning have been introduced into the relation extraction task. In related technologies, common relation extraction models include: PCNN (Piece Wise convolutional neural network) + ONE, PCNN + ATT (attention), PCNN + BAG (bag) - ATT, etc.

[0087] Among them, the PCNN + ONE model uses a piecewise convolutional neural network to extract features, and then selects the most reliable sample from the bag as the representation of the bag according to multi-instance learning. Refer to Figure 1 As shown, taking the sentence “… little Li Si, the son of Zhang San, in…” as an example, the specific implementation process is as follows: First, the input sentence is encoded, and the encoding result is represented by a vector (Vector representation), including word vectors (word) and position vectors (position). After concatenating the word vector and the position vector and extracting features through a convolutional neural network (Convolation), the extracted features pass through a Piecewise max pooling (segmented maximum pooling) layer, and then are concatenated and sent to a Softmax classifier (softmax classifier) layer to finally obtain the classification of the relation. However, this model only utilizes the information of one sentence in the bag and ignores the information of other sentences.

[0088] For the PCNN + ATT model, the representation of each sentence in the bag is obtained by using a piecewise convolutional neural network, and then the attention mechanism is used to assign a weight to each sentence in the bag. The weighted sum of the representations of each sentence is used to obtain the representation of the bag. However, when using the attention mechanism within the bag, the problem of noisy bags cannot be solved, which easily leads to incorrect labeling of the entire bag.

[0089] For the PCNN + BAG - ATT model, it also first uses a piecewise convolutional neural network to learn the representation of each sentence in the bag, and then uses the attention mechanism within the bag to obtain the representation of the bag. In addition, this method also uses the attention mechanism between bags to perform a weighted sum of the representations of the bags to obtain the representation of the group. However, this model is prone to overfitting noisy data and is not robust to noise.

[0090] In summary, the methods in related technologies basically tend to ignore the information of other sentences in the bag, are not robust to noisy data, are prone to overfitting noisy data, and cause performance loss.

[0091] This application provides a method, apparatus, and device for training an information extraction model and obtaining a knowledge graph. A robust training framework based on meta-learning is proposed to train a more robust target information extraction model. Specifically, based on a training sample set containing noisy samples and non-noisy samples, an initial first information extraction model and an initial second information extraction model with the same parameters are trained. First, the initial first information extraction model is trained using the noisy sample set. Based on the update method of meta-learning, the process of the wrong label affecting the model is simulated, and the gradient of the first information extraction model is updated to obtain a first intermediate model. Here, during the process of training the first information extraction model to obtain the first intermediate model, the wrong label is used as the annotation for training. Therefore, the first intermediate model obtained still fits the noisy data. Among the probability values corresponding to multiple labels calculated by the model, the probability value corresponding to the entity association information represented by the noisy label will be relatively high, while the probability value corresponding to the entity association information represented by the corresponding non-noisy label will be relatively low. Then, based on the difference between the prediction result of the noisy samples by the updated first intermediate model and the prediction result of the non-noisy samples by the initial second information extraction model, the initial first information extraction model is updated. At this time, there is a certain difference between the prediction result of the noisy samples by the first intermediate model and the prediction result of the non-noisy samples by the initial second information extraction model. Based on this difference, the parameters of the initial first information extraction model are continuously adjusted to achieve the effect of increasing the prediction probability of the first information extraction model for the correct label and decreasing the prediction probability of the first information extraction model for the wrong label, so that the updated second intermediate model no longer fits the noisy data, ensuring that the updated second intermediate model can give the same prediction result as the initial second information extraction model. Then, the parameters of the second intermediate model are updated using the traditional gradient update method to obtain a trained reference model. Finally, based on the smoothing coefficient, the parameters of the reference model are adjusted to obtain a target information extraction model with better performance. This target information extraction model also does not fit the noisy data and has better robustness. Therefore, the knowledge graph constructed based on this target information extraction model is more accurate. During the above training process, the first information extraction model, the first intermediate model, the second intermediate model, and the reference model all belong to the student models, and the second information extraction model and the finally obtained target information extraction model belong to the teacher models. Finally, the teacher model is used as the trained target information extraction model during the model usage process. Here, the teacher model is obtained through a self-enhancing method and has better prediction ability than the student models.

[0092] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0093] As Figure 2 shown, it is a schematic diagram of the application scenario of the embodiment of the present application. This application scenario diagram includes two terminal devices 210 and a server 220. The terminal devices 210 and the server 220 can communicate through a communication network.

[0094] In an alternative embodiment, the communication network is a wired network or a wireless network. The terminal 210 and the server 220 can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this here.

[0095] In the embodiment of the present application, the terminal device 210 is an electronic device used by a user. This electronic device can be a personal computer, mobile phone, tablet computer, notebook, e-reader, etc., which are computer devices with certain computing capabilities and running instant messaging software and websites or social software and websites. Each terminal device 210 is connected to the server 220 through a wireless network. The server 220 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this here.

[0096] Among them, the target information extraction model can be deployed on the server 220 for training. A large number of training samples, including noise samples and non-noise samples, can be stored in the server 220 for training the target information extraction model. Optionally, after the target information extraction model is trained based on the training method in the embodiment of the present application, the trained target information extraction model can be directly deployed on the server 220 or the terminal device 210. Generally, the target information extraction model is directly deployed on the server 220. In the embodiment of the present application, the target information extraction model is often used to extract entities and the relationships between entities in unstructured natural language texts, and further a knowledge graph, etc. can be constructed based on the entities and entity relationships obtained after processing the natural language text by the target information extraction model.

[0097] It should be noted that the method for training an information extraction model and obtaining a knowledge graph provided in the embodiments of the present application can be applied to scenarios such as advertisement recommendation and library information retrieval. Correspondingly, the training samples used in different scenarios are different. Taking the advertisement recommendation scenario as an example, the training sample used is an advertisement sentence pack, which contains multiple advertisement sentences. Based on this, a knowledge graph related to advertisement recommendation can also be constructed; for example, in the information retrieval scenario, the training sample used is a sentence pack related to literature content, etc. Based on this, a knowledge graph related to literature retrieval can also be constructed.

[0098] In addition, further subdividing, for the advertisement recommendation scenario, it can be specifically divided into advertisement recommendations for different contents such as games, music, pictures, etc. For advertisement recommendations for different contents, the training samples used are also different. Taking game advertisement recommendation as an example, the training sample used is a sentence pack related to games. For example, the sentence pack can contain multiple advertisement sentences related to game content; taking music advertisement recommendation as an example, the training sample used is a sentence pack related to music. For example, the sentence pack can contain multiple advertisement sentences related to music content, etc.

[0099] Similarly, when using the trained target information extraction model to extract entity association information in the text to be processed, in different scenarios, the text to be processed is also different. For example, in the advertisement recommendation scenario, the text to be processed is an advertisement sentence; in the information retrieval scenario, the text to be processed can be literature content, etc.

[0100] In the following, mainly taking the construction of a knowledge graph of character relationships as an example, correspondingly, the noise samples or non-noise samples used are both sentence packs containing character relationships, and the text to be processed is also text containing character relationships.

[0101] Refer to Figure 3 As shown, it is a flowchart of the implementation of a method for training an information extraction model provided in the embodiments of the present application. The specific implementation process of this method is as follows:

[0102] Step S31: Obtain a training sample set including a non-noise sample set and a noise sample set, where each non-noise sample in the non-noise sample set is a sentence pack labeled with a non-noise label, and each noise sample in the noise sample set is a sentence pack labeled with a noise label. Each sentence pack includes multiple sentences for describing the association information between entities, and the entity association information includes the relationship between entities or the events involved by the entities;

[0103] In the embodiments of the present application, the information extraction model can be a relation extraction model for extracting entities and the relations between entities, or an event extraction model for extracting entities and the events involved by the entities. However, both types of models need to first extract entities and then can identify entity relations or event types. The extraction of entities mainly includes the identification of entity locations and entity categories.

[0104] In information extraction, an event generally refers to something that occurs within a specific time segment and geographical scope, involves one or more roles, and consists of one or more actions, usually at the sentence level. The main objective of event extraction is to extract the required interesting event information from text containing event information and present the event expressed in natural language in a structured form. The role of relation extraction is to obtain the syntactic or semantic connections existing between entities in the text, that is, to perform relation classification.

[0105] Correspondingly, when training the information extraction model, the sentences in the sentence package used as training samples refer to unstructured natural language texts, such as sentences in the body of online articles, sentences in the document abstracts in the library database, etc., and these sentences are all used to describe the association information between entities.

[0106] For example, in the sentence "Xiaoming was born in a certain city", "Xiaoming" is a person entity and "a certain city" is a location entity. A relation refers to the semantic relation between entities, such as "friend", "located in", "native place", etc. For example, in the sentence "Xiaoming was born in a certain city", the relation between the two entities "Xiaoming" and "a certain city" is "born in". Another example is the sentence "A concert was held in Place A", where the entity is "Place A" and the event involved is "concert".

[0107] In the embodiments of the present application, the training sample set used for training the information extraction model includes two types of samples. One type is non-noise samples obtained based on the distant supervision method, and the other type is noise samples generated based on the non-noise samples. Among them, both the non-noise samples and the noise samples are sentence packs with labeled tags. The tags labeled on the sentence packs represent the relationships between entities in the sentence packs, such as father-son, mother-son, teacher-student, etc. In addition, each tag also includes a corresponding probability value. The probability value represents the probability that the relationship between the extracted entities belongs to the relationship type represented by the tag. For example, the relationship types include father-son, mother-son, teacher-student, husband-wife, etc. If the relationship type represented by the tag is father-son, the corresponding probability value represents the probability that the relationship between the two entities belongs to father-son; the probability value can also represent the probability that the event involved by the extracted entities belongs to the event type represented by the tag. For example, the event types include concert, sports meeting, parade, etc. If the event type represented by the tag is concert, the corresponding probability value represents the probability that the event involved by the extracted entities belongs to concert.

[0108] In the noise sample, the noise tag is (A, B, father-son relationship, 1). Generally, the tag of father-son relationship is incorrect. However, in the embodiments of the present application, the noise sample is only used to simulate the training process of the model under noise data and is used to train the first intermediate model that will still fit the noise data. Therefore, the corresponding probability can be represented by 1 (for this noise sample, the probability value corresponding to the mother-son relationship is 0). But in fact, the relationship between entity A and entity B is mother-son relationship. In the non-noise sample, the non-noise tag is (A, B, mother-son relationship, 1). Generally, the tag of mother-son relationship is correct. Therefore, the corresponding probability can be represented by 1 (for this non-noise sample, the probability value corresponding to the father-son relationship is 0).

[0109] However, since the non-noise tags are obtained based on the distant supervision method, there will inevitably be cases of incorrect annotation. In the related art, the information extraction model trained only based on the non-noise sample set is very likely to overfit the noise data. In the embodiments of the present application, meta-learning is applied to the distant supervision relationship extraction task, and the model is trained based on the non-noise samples and the noise samples to solve the problem of noise annotation. The target information extraction model trained based on the training sample set in the embodiments of the present application does not fit the noise data and is more robust in the case of noisy data sets.

[0110] In the process of training the target information extraction model based on the training sample data set in step S31, two models are adopted: the first information extraction model and the second information extraction model. The parameters of the initialized first information extraction model and the initialized second information extraction model are the same. Among them, the first information extraction model is used as the student model, and the second information extraction model is used as the teacher model. The second information extraction model is used to guide the training of the first information extraction model. In the embodiment of the present application, these two models are trained by the update method of meta-learning, so that the finally obtained target information extraction model will not fit the noisy data after being trained with noisy data.

[0111] Taking relation extraction as an example, specifically, first, the initial first information extraction model is trained with noisy samples to obtain an updated first intermediate model. This process is mainly to simulate the training process of the initial first information extraction model in the case of noisy data. Since noisy samples with noisy labels (equivalent to wrong labels) are used as training data during the training process, the predicted probability value of the first intermediate model obtained based on this training for the noisy label is higher than the actual probability value, and the predicted probability value for the non-noisy label corresponding to the noisy label is lower than the actual probability value. For example:

[0112] For a non-noisy sample (sentence package B1, mother-child relationship, 1), a corresponding noisy sample is (sentence B1, father-son relationship, 1). In fact, the actual probability value corresponding to the father-son relationship label should be 0. However, when training the initial first information extraction model based on this noisy sample (sentence package B1, father-son relationship, 1), the label is wrong. Among the probability values corresponding to multiple labels predicted by the first intermediate model obtained based on the wrong label training, the probability value corresponding to the father-son relationship is higher, and the probability value corresponding to the mother-child relationship is lower. Finally, it is ensured that the probability value corresponding to the father-son relationship is the largest, that is, the model prediction result indicates that the sentence package B1 represents a father-son relationship. That is to say, the first intermediate model obtained based on this noisy sample training will still fit the noisy data. For example, the probability value corresponding to the father-son relationship predicted by the first intermediate model is 0.65, and the probability value corresponding to the mother-child relationship predicted is 0.2.

[0113] On this basis, based on the difference between the prediction result of the noise samples by the first intermediate state model and the prediction result of the non-noise samples by the initial second information extraction model, continuously adjust the parameters of the initial first information extraction model to achieve the effect of increasing the prediction probability of the first information extraction model for the correct label and decreasing the prediction probability of the first information extraction model for the wrong label, so that the updated second intermediate state model no longer fits the noise data. Therefore, among the probability values corresponding to the multiple labels predicted based on the second intermediate state model, the probability value corresponding to the parent-child relationship will decrease, and the probability value corresponding to the mother-child relationship will increase, ultimately ensuring that the probability value of the mother-child label is the highest. For example, the probability value corresponding to the mother-child relationship predicted based on the second intermediate state model is 0.9, and the probability value corresponding to the parent-child relationship predicted is 0.05.

[0114] Next, adopt the traditional gradient update method to update the second intermediate state model again based on the non-noise samples to obtain the trained reference model. For example, the probability value corresponding to the mother-child relationship predicted based on the reference model is 0.95, and the probability value corresponding to the parent-child relationship predicted is 0.01; finally, adjust the trained reference model through the self-enhancement method to obtain the trained target information extraction model. For example, the probability value corresponding to the mother-child relationship predicted based on the target information extraction model is 0.98, and the probability value corresponding to the parent-child relationship predicted is 0.001.

[0115] The following is a detailed introduction to the above summarized training process. For details, see steps S32 to S35:

[0116] Step S32: Train the initial first information extraction model using the noise sample set, and during the training process, update the first information extraction model based on the difference between the prediction result of the noise samples by the first information extraction model and the noise labels corresponding to the noise samples to obtain the first intermediate state model;

[0117] Step S33: Update the first information extraction model based on the difference between the prediction result of the noise samples by the first intermediate state model and the prediction result of the non-noise samples by the initial second information extraction model to obtain the second intermediate state model, where the parameters of the initial first information extraction model and the initial second information extraction model are the same, and the second information extraction model is used to guide the training of the first information extraction model;

[0118] Step S34: Update the second intermediate state model based on the difference between the prediction result of the non-noise samples by the second intermediate state model and the non-noise labels corresponding to the non-noise samples to obtain the trained reference model;

[0119] Step S35: Adjust the parameters of the trained reference model based on a preset smoothing coefficient to obtain a target information extraction model for obtaining entity association information, where the entity association information is used to construct a knowledge graph.

[0120] When constructing a knowledge graph based on entities and entity relationships, the constructed knowledge graph is mainly used to describe various entity concepts and their mutual relationships. Generally, it consists of triples of "entity-relationship-entity", and each entity has its corresponding "attributes". For example, entities and entity relationships are stored in the form of <subject, relation, object> triples one by one. Subject and object are two entities, and relation represents the relationship between these two entities. Large-scale knowledge graphs often contain hundreds of millions of entities, tens of billions of attributes, and hundreds of billions of relationships, which are mined from a large amount of structured and unstructured data. Based on a dedicated knowledge graph and the natural language understanding technology built on it, machines can give full play to the system performance of reasoning and judgment, answer questions relatively accurately, and extend the scope of intelligence. When constructing a knowledge graph based on entities and the events involved by the entities, the principle is similar, and a knowledge graph describing the relationships between events can be constructed, etc.

[0121] In addition, the knowledge graph constructed based on the information such as entities, relationships, and events that have been obtained can be used for the mutual verification of information and the automatic discovery of abnormal events, such as the combination of multi-level shareholder information and contract information in the legal field to discover related transactions, the missing proof documents of real estate certificates, etc. The task of knowledge graph application is to use the knowledge graph to establish a knowledge-based system and provide intelligent knowledge services, which is the ultimate goal of building a knowledge graph. It mainly includes: information fusion of knowledge-based Internet resources, semantic search, knowledge-based question answering systems, and knowledge-based big data analysis and mining. The knowledge graph not only enables the computer to better understand the knowledge content of Internet resources, but also provides the computer with a better structure for organizing and managing massive data resources.

[0122] It should be noted that the above-listed steps S32 to S35 are the actual training processes when only using a batch of non-noise samples and the noise samples corresponding to this batch of non-noise samples. If there are many training samples, the non-noise samples can be divided into many small batches of samples. For each small batch of samples, the process of steps S32 to S35 needs to be executed to adjust the trained reference model once. When using another small batch of samples for training again, the trained reference model is used as the first information extraction model, and the trained target information extraction model is used as the second information extraction model, and the above process is repeated again, that is, multiple rounds of iterative training are performed.

[0123] The following mainly takes relation extraction as an example for detailed introduction. At this time, the target information extraction model is a relation extraction model, and the sentences in the sentence package are all used to represent the relationships between entities.

[0124] In an alternative embodiment, the set of noise samples is M, where M is a positive integer; when obtaining M sets of noise samples by performing label transfer on the labels of each non-noise sample in the non-noise sample set, each noise sample in the set of noise samples can be obtained through the following method:

[0125] For each non-noise sample in the non-noise sample set, perform M times of label transfer on each non-noise sample respectively according to the non-noise labels of other non-noise samples except itself to obtain M noise labels corresponding to each non-noise sample; then, based on the sentence package included in each non-noise sample and the M noise labels corresponding to each non-noise sample, combine and generate M noise samples corresponding to each non-noise sample.

[0126] For example, a non-noise sample 1 is a sentence package B1 labeled with a father-son relationship, with the label: father-son relationship. The sentence package B1 contains multiple sentences that simultaneously contain entity A and entity B, such as A and B are father and son, B is A's father, A is taller than B, and the household registrations of A and B are both in a certain city, etc. And the non-noise sample 2 is a sentence package B2 labeled with a mother-son relationship, with the label: mother-son relationship. The sentence package B2 contains multiple sentences that simultaneously contain entity A and entity C. After transferring the label of the sentence package B2 to the sentence package B1, a sentence package B1 labeled with a mother-son relationship is obtained, that is, a noise sample corresponding to the non-noise sample 1.

[0127] Based on the above method, finally each non-noise sample will correspond to M noise samples, and the minimum value of M is 1. Taking M = 3 as an example, for the non-noise sample B 1 (sentence package), the non-noise label of this sentence package is y 1 , then after performing 3 times of label transfer on this non-noise sample, 3 corresponding noise samples can be obtained. These 3 noise samples are all the sentence package B1, and the difference is that the noise labels corresponding to these 3 noise samples are different, which are respectively

[0128] Table 1

[0129]

[0130] Table 1 is a schematic diagram showing the results of three label transfers for three non-noise samples in an embodiment of the present application. Among them, for the non-noise sample (B1, parent-child relationship), the noise sample 1 obtained after the first label transfer is (B1, mother-child relationship), the noise sample 1 obtained after the second label transfer is (B1, teacher-student relationship), and the noise sample 1 obtained after the third label transfer is (B1, friend relationship). The non-noise sample and the corresponding three noise samples contain the same sentence package, which is B1, but only the labeled tags are different. Among them, the parent-child relationship belongs to the non-noise label, while the mother-child relationship, teacher-student relationship, and friend relationship are the corresponding noise labels. The same applies to the non-noise sample (B2, mother-child relationship), and the corresponding three noise samples are (B2, teacher-student relationship), (B2, friend relationship), and (B2, lover relationship), and the sentence packages contained are all B2. For the non-noise sample (B3, teacher-student relationship), the corresponding three noise samples are (B3, friend relationship), (B3, parent-child relationship), and (B3, spousal relationship), and the sentence packages contained are all B3. Among them, the noise samples in the same row in Table 1 are obtained from the same label transfer and belong to the same noise sample set, and the noise samples in different rows belong to different noise sample sets.

[0131] Through the above process, the actually obtained noise set will contain M batches of noise samples, that is, obtained by performing M label transfers on each non-noise sample in the non-noise set. These noise samples in the above process are generated by imitating the distribution of the original non-noise samples.

[0132] The generation process of each noise sample in the noise sample set will be introduced in detail below.

[0133] Suppose that in each training, a small batch of data (X, Y) is first sampled from the non-noise sample set. Among them, X = {B 1 , B 2 , …, B k}, which contains k sentence packages. Each sentence package is composed of multiple sentences describing the relationship between entities, and each sentence package corresponds to a non-noise label. Y = {y 1 , y 2 , …, y k} contains k labels, which represent the non-noise labels corresponding to each sentence package. For X, M batches of noise labels can be generated through the above method, denoted as Among them That is, the noise label of the sentence package B k obtained in the m-th label transfer. The process of label transfer will be introduced in detail below taking as an example, and mainly two methods are listed:

[0134] Method 1: Label transfer based on random selection.

[0135] The specific implementation method is that for any non-noise sample, the non-noise label of any other non-noise sample except this non-noise sample is used as the noise label obtained after label transfer for this non-noise sample.

[0136] For example, first randomly select 5 sentence packs from this small batch of data (X, Y). For any sentence pack B among the randomly selected 5 sentence packs i , when obtaining a noise label by label transfer for this sentence pack, another sentence pack B can be randomly selected from the remaining 4 sentence packs except B i , and its non-noise label y j is used as a noise label for sentence pack B j , that is i , where i ≠ j. where i ≠ j.

[0137] Method 2: Label transfer based on similarity.

[0138] The specific implementation method is that for any non-noise sample, the similarity between this non-noise sample and other non-noise samples except non-noise samples is obtained. Among the top N other non-noise samples sorted by similarity from high to low, the non-noise label of any other non-noise sample is used as the noise label obtained after label transfer for this non-noise sample, where N is a positive integer.

[0139] For example, first randomly select 5 sentence packs from this small batch of data (X, Y). For any sentence pack B among the randomly selected 5 sentence packs i , first calculate its similarity with the other 4 sentence packs, then sort based on the similarity, and finally select a sentence pack B from the N sentence packs closest to sentence pack B i , and use its non-noise label to replace the non-noise label of sentence pack B j , then a noise label for sentence pack B i can be obtained, that is i , where i ≠ j. where i ≠ j.

[0140] In the above implementation, after obtaining the noise labels corresponding to each non-noise sample, the sentence packs included in each non-noise sample are combined with the noise labels to obtain sentence packs labeled with noise labels, that is, noise samples. Since the noise labels generated in the above method come from the neighbors of non-noise samples, the noise samples constructed based on similarity follow a similar distribution to the original non-noise samples.

[0141] In addition, it should be noted that when randomly selecting multiple sentence packs from a small batch of data (X, Y), the idea in Method 2 can also be combined to directly select multiple sentence packs that are relatively close, so that there is a neighbor relationship between these sentence packs. Then, when performing label transfer, for the selected sentence pack B i , use the label of its neighbor to replace the label of sentence pack B i . The noisy labels generated in this way all come from neighbors. Therefore, the constructed noisy samples and the original non-noisy samples also have a similar distribution. The model trained based on such training samples can more accurately analyze the differences between similar samples and has higher robustness.

[0142] Among them, the similarity between sentence packs means that the representations of the sentence packs are relatively close. In the embodiments of the present application, the representation of a sentence pack can be in the form of a vector. Therefore, when determining the similarity between sentence packs, the Euclidean distance or cosine distance can be used. First, encode the sentence pack into a vector form, and then calculate the Euclidean distance or cosine distance between the vectors, etc., to determine the similarity between the sentence packs, that is, the similarity between non-noisy samples. The closer the distance, the higher the similarity between non-noisy samples.

[0143] For example, when N = 3, 7 sentence packs are randomly selected from the data (X, Y), which are B 1 , B 2 , …, B 7 . For B 1 , among the remaining 6 sentence packs, the top 3 sentence packs with the highest similarity to it are B 2 , B 3 , B 4 . At this time, if sentence pack B 2 , B 3 , B 4 are randomly selected to get sentence pack B 2 , and the non-noisy label of sentence pack B 2 is used as a noisy label of sentence pack B 1 , then we can get Assume that the next time sentence pack B 3 is randomly selected, and the non-noisy label of sentence pack B 3 is used as a noisy label of sentence pack B 1 , then we can get

[0144] In the embodiments of the present application, in order to better model label noise, the above process is repeated M times, and then M small batches of artificially generated noisy labels can be obtained, denoted as Among them represents a noisy label corresponding to each non-noisy sample in this batch of data.

[0145] In the above embodiment, by imitating the distribution of the original non-noisy samples, multiple batches of noisy samples are generated. When the number of samples increases, the accuracy of the trained model is also higher.

[0146] After introducing the training sample set, the following will introduce in detail the model training process in the embodiments of the present application in combination with Figure 4 Refer to

[0147] As shown in Figure 4 , which is a robustness training framework for a remote supervision relation extraction model based on meta-learning in the embodiments of the present application. This framework mainly includes the following modules:

[0148] (1) Generation of Synthetic Noisy Labels (GSNL) module: Imitating the distribution of the original non-noisy samples, generating some noisy labels, constructing noisy samples, and using them for the gradient update of meta-learning;

[0149] Among them, (X, Y) is a small batch of non-noisy samples, and is a small batch of noisy samples obtained by generating fitting noisy labels. Based on this module, the noisy sample set in step S31 can be generated. The generation process of the fitting noisy labels can refer to the label transfer methods based on random selection, similarity-based label transfer methods, etc. listed in the above embodiments, which will not be elaborated here.

[0150] (2) Meta-Train module: Using the generated artificial noisy samples to simulate the gradient update process of the relation extraction model under noisy data;

[0151] For example, using PCNN+ATT as a relation extraction model, called the first information extraction model (student model), and the parameters of the model are θ, denoted as f(θ). Based on the first information extraction model, the second information extraction model (teacher model) can be constructed by using the self-ensemble method, denoted as

[0152] Among them, the parameters of the initial first information extraction model and the initial second information extraction model are the same, that is, the initial The initial θ is obtained by random initialization. θ′ 1 refers to a first intermediate model obtained by performing gradient update on the initial first information extraction model based on the non-noisy sample .

[0153] In the embodiment of the present application, the first intermediate state model is obtained by performing one-step gradient update on the initial first information extraction model based on the noise samples obtained through the above process, aiming to simulate the training process of the relation extraction model in the presence of noisy data. Since M batches of noise samples can actually be obtained in the above process, which is expressed as Therefore, when training the initial first information extraction model based on the M batches of noise samples in the noise sample set, it can be carried out batch by batch.

[0154] In an alternative embodiment, the prediction result of the initial first information extraction model for the noise samples is the first prediction label, that is, the label obtained through the prediction of the first information extraction model; in step S32, when training the initial first information extraction model with the noise sample set and updating the initial first information extraction model based on the difference between the prediction result of the initial first information extraction model for the noise samples and the noise label corresponding to the noise samples during the training process, M first intermediate state models can be obtained. The specific implementation method is as follows:

[0155] Input the noise samples in the M noise sample sets into the initial first information extraction model batch by batch, and obtain M batches of first prediction labels output by the initial first information extraction model. Among them, the noise samples in the same batch belong to the same noise sample set, and the noise samples in different batches belong to different noise sample sets, that is, the M noise sample sets are used as M batches of sample data and input into the initial first information extraction model batch by batch. Based on the M batches of sample data, M batches of prediction labels can be obtained; obtain the first classification loss function determined based on the difference between each batch of first prediction labels and the noise labels of the corresponding noise samples; respectively perform one-step gradient update on the initial first information extraction model according to each first classification loss function to obtain M first intermediate state models.

[0156] Among them, the difference between the first prediction label and the noise label of the corresponding noise sample is determined based on the probability value. When the model outputs the first prediction label, the corresponding probability value is also output, that is, the probability that the entity relationship represented by the sentence pack predicted by the model is the relationship type represented by the first prediction label.

[0157] For example, the first prediction label indicates that the probability of the relationship between entity 1 and entity 2 in the sentence pack being parent-child is 0.85, the corresponding noise label indicates that the relationship between entity 1 and entity 2 is parent-child, and the corresponding probability value is 1. Then, based on the difference between these two probability values, the first classification loss function is calculated, and further, one-step gradient update is performed on the initial first information extraction model.

[0158] Specifically, for any batch of noise samples After inputting the noise samples into the initial first information extraction model, based on the difference between the first predicted label output by the initial first information extraction model and the noise label corresponding to the noise sample, perform one-step gradient update on the initial first information extraction model, changing from θ to θ′ m , that is, the first intermediate model, where θ′ m can be expressed by the following formula:

[0159]

[0160] where, represents the first classification loss function determined based on the difference between the first predicted label output by the initial first information extraction model and the corresponding noise label . In the embodiments of the present application, the cross-entropy loss function is adopted, and α represents the step size, generally taking a value of about 0.01.

[0161] Since one θ′ m can be obtained based on each batch of noise samples, after inputting the noise samples into the initial first information extraction model in batches, M first intermediate models θ′ m can be obtained.

[0162] In the embodiments of the present application, the initial student model can be a neural network model such as PCNN+ATT, PCNN+BAG-ATT, etc. that can be used for relation extraction. Among them, PCNN is a segmented convolutional neural network for encoding sentence information, and here PCNN can also be replaced with other models, such as LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), Transformer, etc.

[0163] Taking the PCNN+ATT model as an example, when using this model to perform relation extraction on the sentence bag B = {x 1 , x 2 , x 3 , …, x n}, the specific implementation process is as shown in Figure 5A . First, represent each sentence x in the bag as a vector form x through the segmented convolutional neural network to obtain the representation of each sentence in the bag, and then use the attention mechanism to assign a weight α to each sentence in the bag, and perform weighted summation on the representation of each sentence based on α to obtain the representation s of the bag.

[0164] Taking the PCNN+BAG-ATT model as an example, when using this model to perform relation extraction on the sentence bag, it also first uses the segmented convolutional neural network to learn the representation of each sentence in the bag, as shown in Figure 5B . Among them, the input sentence is encoded by the sentence encoder Encode to obtain the representation of each sentence Here it is in the form of a vector. Then, combined with the relation matrix R, use the intra-bag attention mechanism to obtain the representation B of the bag n , and then use the inter-bag attention mechanism to perform weighted summation on the representations of the bags to obtain the representation G of the group. Both the representation of the bag and the representation of the group are in the form of matrices.

[0165] (3) Meta-Test module: The purpose is to enable the relation extraction model not to overfit the noisy data after training on the noisy data, that is, train θ so that the updated θ′ m does not overfit the noise. To achieve this goal, in the embodiments of the present application, it is required that the updated student model (the first intermediate model) gives the same prediction as the teacher model. For this purpose, in the embodiments of the present application, a consistency loss function based on the KL (Kullback-Leibler) divergence (alias: information divergence / relative entropy) is proposed This function is used to measure the difference between the prediction of the updated student model and the prediction result of the teacher model.

[0166] In an optional implementation manner, the prediction result of the first intermediate model for the noise samples is the second prediction label, and the prediction result of the initial second information extraction model for the non-noise samples is the third prediction label; in step S33, based on the difference between the prediction result of the first intermediate model for the noise samples and the prediction result of the initial second information extraction model for the non-noise samples, when updating the initial first information extraction model to obtain the second intermediate model, the specific implementation manner is as follows:

[0167] Input the noise samples in the M noise sample sets into the corresponding first intermediate models respectively to obtain the second prediction labels output by each first intermediate model; and input the non-noise samples in the non-noise sample set into the initial second information extraction model to obtain the third prediction labels output by the initial second information extraction model; determine the consistency loss function based on the KL divergence respectively according to the difference between each batch of second prediction labels and third prediction labels; average the M consistency loss functions to obtain the fused difference; perform at least one gradient update on the initial first information extraction model according to the fused difference to obtain the second intermediate model, where the error between the prediction result of the second intermediate model and the prediction result of the initial second information extraction model is within the specified range.

[0168] Since the noise samples are divided into M batches in the above process, M first intermediate models θ′ are obtained m, therefore, when performing meta-testing based on this, it is necessary to input M batches of noise samples into the corresponding first intermediate state model θ' respectively m .

[0169] Among them, the first intermediate state model corresponding to the noise sample refers to θ' m which is corresponding to , for example, for the first batch of noise samples the first intermediate state model θ' trained based on this batch of noise samples 1 is the first intermediate state model corresponding to this batch of noise samples. Similarly, for the second batch of noise samples the corresponding first intermediate state model is θ' 2 , …, and so on. When obtaining the second predicted label output by the first intermediate state model, it is only necessary to input M batches of noise samples into the corresponding first intermediate state model respectively. The second predicted label output by the first intermediate state model can be expressed as f(X, θ' m ), which represents the probability that the entity relationship represented by the sentence pack predicted by the model is the relationship type represented by the second predicted label.

[0170] In addition, it is also necessary to input the non-noise samples (X, Y) into the second information extraction model to obtain the third predicted label predicted for each non-noise sample output by the second information extraction model which represents the probability that the entity relationship represented by the sentence pack predicted by the model is the relationship type represented by the third predicted label.

[0171] For example, the second predicted label indicates that the probability that the relationship between entity 1 and entity 2 in the sentence pack is father-son is 0.65 (the probability value corresponding to the predicted mother-son relationship is 0.2), and the corresponding third predicted label indicates that the probability that the relationship between entity 1 and entity 2 is father-son is 0.05 (the probability value corresponding to the predicted mother-son relationship is 0.9). Then, based on the difference between these two probability values of 0.65 and 0.05, calculate the consistency loss function. After constructing the meta-objective function based on the consistency loss function, continuously adjust the first intermediate state model based on the meta-objective loss function to reduce the difference between the second predicted label and the third predicted label, so that the second predicted label is close to the third predicted label, and obtain the updated second intermediate adjustment model. The meta-objective loss function in the embodiments of the present application is the fusion difference. When M = 1, there is no need to calculate the average of the consistency loss function to calculate the fusion difference, and the consistency loss function can be directly calculated.

[0172] In the embodiments of the present application, the consistency loss function based on KL divergence has the following specific calculation formula:

[0173]

[0174] Since M batches of noise samples are constructed in the above process, it is necessary to minimize the consistency loss function for M first intermediate state models θ′ m to obtain the meta-objective function. The meta-objective function is obtained by averaging M consistency loss functions and can be expressed as defined as follows:

[0175]

[0176] where

[0177] In the embodiment of the present application, based on the constraint of the meta-objective function, the first information extraction model (student model) is updated, and the calculation formula is as follows:

[0178]

[0179] where η represents the learning efficiency of meta-learning, and its value range is generally 10 -3 . By minimizing the meta-objective function through the gradient descent method, based on the gradient of the meta-objective function, the initial first information extraction model is updated, and the model obtained when infinitely approaches θ is used as the second intermediate state model, that is, the updated student model, and the parameters are still represented by θ.

[0180] In the above implementation, through the constraint of the meta-objective function, the updated student model (second intermediate state model) can give the same prediction result as the teacher model, thus ensuring that the updated student model does not overfit the noise data.

[0181] In an alternative implementation, the prediction result of the second intermediate state model for non-noise samples is the fourth prediction label; in step S34, based on the difference between the prediction result of the second intermediate state model for non-noise samples and the non-noise label corresponding to the non-noise sample, when updating the second intermediate state model to obtain the trained reference model, the specific process is as follows:

[0182] Input the non-noise samples in the non-noise sample set into the second intermediate state model to obtain the fourth prediction label output by the second intermediate state model; obtain the second classification loss function determined based on the difference between the fourth prediction label and the non-noise label of the corresponding non-noise sample; according to the second classification loss function, perform at least one gradient update on the second intermediate state model to obtain the trained reference model, where the error between the prediction result of the trained reference model and the prediction result of the second intermediate state model is within the specified range.

[0183] This process is the process of performing gradient update on the second intermediate tone model using the original non-noisy samples. Since after the meta-learning update, the obtained second intermediate state model does not fit the noisy data, therefore, on this basis, the second intermediate state model can be updated by minimizing the second classification loss function calculated using the original non-noisy samples (X, Y) in the way of stochastic gradient descent to obtain the trained reference model.

[0184] Among them, the second classification loss function is determined by the difference between the fourth predicted label obtained by inputting the original non-noisy samples (X, Y) into the second intermediate state model and the non-noisy label corresponding to the non-noisy samples, and can be expressed as

[0185] Among them, the fourth predicted label represents the probability that the entity relationship represented by the sentence pack predicted by the model is the relationship type represented by the fourth predicted label.

[0186] For example, the fourth predicted label represents that the probability of the relationship between entity 1 and entity 2 in the sentence pack being father-son is 0.01, and the corresponding non-noisy label represents that the probability of the relationship between entity 1 and entity 2 being father-son is 0 (the non-noisy label actually represents that the probability of the relationship between entity 1 and entity 2 being mother-son is 1, so it can be known that the probability of the relationship between entity 1 and entity 2 being father-son is 0). At this time, the fourth predicted label is already very close to the corresponding non-noisy label. Then, based on the difference between the two probability values of 0.01 and 0, the second classification loss function is calculated, and the second intermediate state model is continuously adjusted based on the second classification loss function to reduce the difference between the fourth predicted label and the corresponding non-noisy label to obtain the trained reference model.

[0187] It can also be expressed as: the fourth predicted label represents that the probability of the relationship between entity 1 and entity 2 in the sentence pack being mother-son is 0.95, and the corresponding non-noisy label represents that the probability of the relationship between entity 1 and entity 2 being mother-son is 1. Then, based on the difference between the two probability values of 0.95 and 1, the second classification loss function is calculated, and the second intermediate state model is continuously adjusted based on the second classification loss function to reduce the difference between the fourth predicted label and the corresponding non-noisy label to obtain the trained reference model. A similar method can also be used when calculating other loss functions.

[0188] In the embodiments of the present application, the process of performing gradient update on the second intermediate state model based on the second classification loss function can be expressed by the following formula:

[0189]

[0190] Among them, β represents the learning efficiency, and its value range is generally 10 -3 , by taking the gradient of the second classification loss function, adjusting the parameter θ of the second intermediate state model, and The model obtained when approaching θ infinitely is used as the trained reference model.

[0191] It should be noted that both the first classification loss function and the second classification loss function in the embodiments of the present application can be represented by the cross - entropy loss function. The cross - entropy loss function is often used to measure the difference between the predicted result distribution and the true result distribution. Compared with other types of loss functions, using cross - entropy as the loss function can avoid gradient dissipation and ensure the learning rate.

[0192] In step S35, when adjusting the parameters of the trained reference model based on the preset smoothing coefficient, it can be done in an exponential averaging manner, that is, the parameters of the trained reference model are exponentially averaged according to the preset smoothing coefficient to obtain the relation extraction model, which can be specifically calculated based on the following formula:

[0193]

[0194] Among them, γ represents the smoothing coefficient, and its value range is 0 to 1. In the embodiments of the present application, generally, γ takes a value near 0.999.

[0195] In the above - mentioned implementation manner, the teacher model obtained through this self - enhancing method generally has better prediction ability than the student model. Therefore, the accuracy of the relation extraction model obtained by adjusting the parameters of the trained reference model based on the preset smoothing coefficient is higher.

[0196] In the embodiments of the present application, any one of the methods for training the information extraction model listed in the above - mentioned embodiments can still be used to train the event extraction model that does not fit the noise data, and the training method is the same. For example, in combination with Figure 4 As shown in the architecture diagram, first, artificial noise samples are constructed through the fitting noise label generation module, and then meta - training and meta - testing are carried out in combination with the teacher model and the student model, and the event extraction model that does not fit the noise data is trained through meta - learning. The specific implementation manner can refer to the above - mentioned embodiments and will not be repeated here.

[0197] To prove the effectiveness of the present application, it was verified on the relevant remotely - supervised relation extraction dataset, and the experimental results are as follows:

[0198] Refer to Figure 6 As shown, it is a schematic diagram of an experimental result in the embodiments of the present application. Figure 6 The curve shown in it is the PR curve of each model, where P refers to precision and R refers to recall. In Figure 6In it, with recall as the horizontal axis and precision as the vertical axis, the comparison results between the present application and the following related models are shown:

[0199] (1) Mintz (a method named after a person): A multi-class logistic regression model using syntactic and lexical features;

[0200] (2) MultiR (a probabilistic graphical model considering multiple relationships): A probabilistic graphical model for multi-instance learning;

[0201] (3) MIMLRE (a relation extraction model based on multi-instance multi-label): A graph model for multi-instance multi-label learning;

[0202] (4) PCNN+ONE (a piecewise convolutional neural network trained by selecting the most reliable samples): A convolutional neural network model using piecewise max pooling;

[0203] (5) PCNN+ATT (a piecewise convolutional neural network combined with sentence-level attention mechanism): A convolutional neural network model based on attention mechanism;

[0204] (6) PCNN+ATT+SL (a piecewise convolutional neural network combined with attention mechanism and soft labels): A method using attention mechanism and soft labels;

[0205] (7) BGWA (a bidirectional gated recurrent unit network combined at the word level): A bidirectional GRU model based on lexical-level and sentence-level attention mechanisms;

[0206] (8) RESIDE (a relation extraction model using additional information): A model using external knowledge and encoding syntactic structures using graph convolutional neural networks;

[0207] (9) PCNN+BAG-ATT (a piecewise convolutional neural network combined with bag-level attention mechanism): A model based on intra-bag attention and inter-bag attention mechanisms;

[0208] (10) TFML (a training framework based on meta-learning), a relation extraction model obtained based on the training method in the present application.

[0209] It can be seen that the results of the model TFML in the present application on the relevant data sets are better than those of the models in the related technologies. As Figure 6 shown, the PR curve corresponding to the TFML model is above that of all the models in the related technologies.

[0210] Table 1: Comparison of the present application and the models in the related technologies in terms of P@N metric

[0211]

[0212] Among them, P@N is the accuracy rate of the first N sentence packs in the training samples. One means randomly taking one sentence from the pack for testing, Two means taking two sentences from the pack for testing, All means using all sentences for testing, and Mean is the average of 100, 200, and 300. Table 1 discusses the comparison of the target information extraction model in this application with the models in related technologies in terms of the P@N metric. It can be seen that the performance of this application is much better than that of the models in related technologies, which also illustrates the effectiveness of the method proposed in this application.

[0213] Table 2: Performance of the models in this application and related technologies under more noisy data

[0214] Method (Model) r=0 r=0.1 r=0.2 r=0.3 r=0.4 PCNN+ONE 68.7 64.4(-4.3) 60.8(-7.9) 58.1(-10.6) 55.9(-12.8) PCNN+ATT 72.2 68.1(-4.1) 64.7(-7.5) 62.0(-10.2) 60.1(-12.1) BGWA 76.3 72.4(-3.9) 69.1(-7.2) 66.5(-9.8) 64.4(-11.9) PCNN+BAG-ATT 84.8 81.3(-3.5) 78.2(-6.6) 76.0(-8.8) 74.3(-10.5) TMFL 84.0 81.9(-2.1) 80.0(-4.0) 78.3(-5.7) 76.6(-7.4)

[0215] Among them, Table 2 discusses the performance of the relation extraction model in this application and the models in related technologies under more noisy data. r represents the proportion of artificially adding noise to the data. It can be seen from the above table that compared with the models in related technologies, the relation extraction model trained by using the method of this application has better performance and is less affected by noisy data under the same noise conditions. When a larger proportion of noise is added to the data, the performance of this application decreases less, which illustrates the effectiveness of the method proposed in this application.

[0216] Refer to Figure 7A As shown, it is the implementation flowchart of a method for obtaining a knowledge graph provided by an embodiment of this application. The specific implementation process of this method is as follows:

[0217] S71: Obtain the text to be processed, where the text to be processed is an unstructured natural language text used to describe the association information between entities, and the entity association information includes the relationship between entities or the events involved by entities;

[0218] S72: Input the text to be processed into the trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model;

[0219] Among them, the target information extraction model is trained by any of the above methods for training the target information extraction model, such as Figure 3 the method shown.

[0220] S73: Construct a knowledge graph based on the entity association information.

[0221] When using the target information extraction model trained by the above method, taking the target information extraction model as a relation extraction model as an example, specifically as Figure 7BAs shown, by inputting unstructured natural language text into a trained relation extraction model, the relation types expressed by entities in the text can be identified and represented in the form of triples <subject, relation, object>. Based on the triples, a knowledge graph can be constructed. Compared with the methods in related technologies, the embodiments of the present application can prevent the parameters of the relation extraction model from fitting noise data. Therefore, the entities and entity relations obtained based on this model are more accurate, and the constructed knowledge graph is also more precise.

[0222] In addition, the target information extraction model can also be an event extraction model, which is mainly applied to event extraction tasks, extracting events of interest to users from the text to be processed describing event information and presenting them in a structured form. Specifically, when performing event extraction, first, the event and its type are identified. Second, the elements involved in the event (usually entities) are identified. Finally, the role played by each element in the event needs to be determined.

[0223] It should be noted that the process of constructing a knowledge graph based on entity association information can refer to the above embodiments, and the repeated parts will not be elaborated.

[0224] Refer to Figure 8 As shown, it is a complete method flow chart for training an information extraction model. The specific implementation process of this method is as follows:

[0225] Step S801: Randomly initialize the student model (model parameters are θ);

[0226] Step S802: Initialize the teacher model (model parameters );

[0227] Step S803: Select a small batch of non-noise samples from the non-noise sample set;

[0228] Step S804: Perform label transfer on the selected small batch of non-noise samples to generate a corresponding batch of noise samples;

[0229] Step S805: Train the initial teacher model based on the generated noise samples, and during the training process, update the initial student model based on the difference between the prediction result of the initial student model for the noise samples and the noise labels corresponding to the noise samples, to obtain the first intermediate model (model parameters are θ' m );

[0230] Step S806: Construct a consistency loss function based on the difference between the prediction result of the first intermediate model for the noise samples and the prediction result of the initial teacher model for the non-noise samples;

[0231] Among them, the consistency loss function is:

[0232]

[0233] Step S807: Determine whether the number of label transfers reaches M times. If so, execute Step S808; otherwise, return to Step S804;

[0234] Step S808: Construct a meta-objective function based on the obtained M consistency loss functions;

[0235] Among them, the meta-objective function is:

[0236]

[0237] Step S809: Perform gradient update on the student model according to the meta-objective function to obtain a second intermediate state model; among them, the calculation formula during gradient update is:

[0238]

[0239] Step S810: Perform gradient update on the second intermediate state model based on the difference between the prediction result of the non-noise sample by the second intermediate state model and the non-noise label corresponding to the non-noise sample to obtain a trained reference model;

[0240] Among them, the calculation formula during gradient update is:

[0241]

[0242] Step S811: Adjust the parameters of the reference model based on a preset smoothing coefficient to obtain an updated teacher model;

[0243] Among them, the calculation formula for adjusting the parameters is:

[0244]

[0245] Step S812: Determine whether there are still non-noise samples in the non-noise sample set. If so, return to Step S803; otherwise, execute Step S813;

[0246] Step S813: Use the teacher model after the last update as the trained target information extraction model.

[0247] It should be noted that in the above process, the initial student model is the initial first information extraction model, and the initial teacher model is the initial second information extraction model.

[0248] In the embodiment of the present application, finally (the teacher model) is used for testing to perform relationship extraction on natural language texts.

[0249] Based on the same inventive concept, an embodiment of the present application further provides an apparatus for training an information extraction model, as Figure 9 shown, which is a schematic structural diagram of the apparatus 900 for training an information extraction model, and may include:

[0250] An acquisition unit 901, configured to acquire a training sample set including a non-noise sample set and a noise sample set, where each non-noise sample in the non-noise sample set is a sentence packet with a labeled non-noise label, each noise sample in the noise sample set is a sentence packet with a labeled noise label, each sentence packet includes multiple sentences for describing the association information between entities, the sentences are unstructured natural language texts, and the entity association information includes the relationship between entities or the events involved by the entities;

[0251] A first update unit 902, configured to train an initial first information extraction model using the noise sample set, and during the training process, update the initial first information extraction model based on the difference between the prediction result of the noise sample by the initial first information extraction model and the noise label corresponding to the noise sample, to obtain a first intermediate model;

[0252] A second update unit 903, configured to update the initial first information extraction model based on the difference between the prediction result of the noise sample by the first intermediate model and the prediction result of the non-noise sample by the initial second information extraction model, to obtain a second intermediate model, where the parameters of the initial first information extraction model and the initial second information extraction model are the same, and the second information extraction model is used to guide the training of the first information extraction model;

[0253] A third update unit 904, configured to update the second intermediate model based on the difference between the prediction result of the non-noise sample by the second intermediate model and the non-noise label corresponding to the non-noise sample, to obtain a trained reference model;

[0254] An adjustment unit 905, configured to adjust the parameters of the trained reference model based on a preset smoothing coefficient, to obtain a target information extraction model for obtaining entity association information, where the entity association information is used to construct a knowledge graph.

[0255] Optionally, the noise sample set is obtained by performing label transfer on the labels of each non-noise sample in the non-noise sample set.

[0256] Optionally, there are M noise sample sets, where M is a positive integer, and the acquisition unit 901 is specifically configured to:

[0257] For each non-noise sample in the non-noise sample set, according to the non-noise labels of other non-noise samples except itself, perform M times of label transfer on each non-noise sample respectively to obtain M noise labels corresponding to each non-noise sample;

[0258] Based on the sentence packs included in each non-noise sample and the M noise labels corresponding to each non-noise sample respectively, generate M noise samples corresponding to each non-noise sample. The sentence packs of the noise samples corresponding to each non-noise sample are the same, and the noise samples obtained by each non-noise sample for each label transfer belong to a noise sample set.

[0259] Optionally, the obtaining unit 901 is specifically configured to:

[0260] For any non-noise sample, the following process is executed every time a corresponding noise label is obtained by label transfer:

[0261] Obtain the similarity between the non-noise sample and other non-noise samples except the non-noise sample, and use the non-noise label of any one of the top N other non-noise samples with the similarity sorted from high to low as the noise label obtained after label transfer of the non-noise sample, where N is a positive integer.

[0262] Optionally, the prediction result of the initial first information extraction model for the noise sample is the first prediction label; the first updating unit 902 is specifically configured to:

[0263] Input the noise samples in the M noise sample sets in batches into the initial first information extraction model to obtain M batches of first prediction labels output by the initial first information extraction model, where the noise samples in the same batch belong to the same noise sample set, and the noise samples in different batches belong to different noise sample sets;

[0264] Obtain the first classification loss function determined based on the difference between each batch of first prediction labels and the noise labels of the corresponding noise samples;

[0265] Perform a gradient update on the initial first information extraction model respectively according to each first classification loss function to obtain M first intermediate models;

[0266] The second updating model 903 is specifically configured to: obtain the fusion difference based on the difference between the prediction results of each first intermediate model for the noise samples and the prediction results of the initial second information extraction model for the non-noise samples, and update the initial first information extraction model according to the fusion difference to obtain the second intermediate model.

[0267] Optionally, the prediction result of the first intermediate state model for the noise samples is the second prediction label, and the prediction result of the initial second information extraction model for the non-noise samples is the third prediction label; the second update unit 903 is specifically configured to:

[0268] Input the noise samples in the M noise sample sets into the corresponding first intermediate state models respectively, to obtain the second prediction labels output by each first intermediate state model; and input the non-noise samples in the non-noise sample set into the initial second information extraction model, to obtain the third prediction labels output by the initial second information extraction model;

[0269] Determine the consistency loss function based on the KL divergence respectively according to the differences between each batch of second prediction labels and third prediction labels;

[0270] Average the M consistency loss functions to obtain the fusion difference;

[0271] Perform at least one gradient update on the initial first information extraction model according to the fusion difference to obtain a second intermediate state model, where the error between the prediction result of the second intermediate state model and the prediction result of the initial second information extraction model is within a specified range.

[0272] Optionally, the prediction result of the second intermediate state model for the non-noise samples is the fourth prediction label; the third update unit 904 is specifically configured to:

[0273] Input the non-noise samples in the non-noise sample set into the second intermediate state model to obtain the fourth prediction labels output by the second intermediate state model;

[0274] Obtain the second classification loss function determined based on the difference between the fourth prediction label and the non-noise label of the corresponding non-noise sample;

[0275] Perform at least one gradient update on the second intermediate state model according to the second classification loss function to obtain a trained reference model, where the error between the prediction result of the trained reference model and the prediction result of the second intermediate state model is within a specified range.

[0276] Optionally, the adjustment unit 905 is specifically configured to:

[0277] Perform exponential averaging on the parameters of the trained reference model according to a preset smoothing coefficient to obtain a target information extraction model.

[0278] Based on the same inventive concept, an embodiment of the present application further provides a device for obtaining a knowledge graph, as Figure 10 shown, which is a schematic structural diagram of the device 1000 for obtaining a knowledge graph, and may include:

[0279] An acquisition unit 1001, configured to acquire a text to be processed, where the text to be processed is an unstructured natural language text for describing the association information between entities, and the entity association information includes the relationship between entities or the events involved by entities;

[0280] An information extraction unit 1002, configured to input the text to be processed into a trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model, where the target information extraction model is obtained by training using any of the above methods for training the target information extraction model;

[0281] A construction unit 1003, configured to construct a knowledge graph based on the entity association information.

[0282] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.

[0283] Those skilled in the art of the technical field to which this application pertains can understand that various aspects of this application can be implemented as a system, a method, or a program product. Therefore, various aspects of this application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0284] In some possible implementation manners, the embodiment of this application further provides an electronic device. Referring to Figure 9 as shown, the electronic device 1100 may at least include a processor 1101 and a memory 1102. Among them, the memory 1102 stores program code. When the program code is executed by the processor 1101, the processor 1101 is caused to execute the steps in the content recommendation method according to various exemplary implementation manners of this application described above in this specification. For example, the processor 1101 may execute the steps as shown in Figure 3 or Figure 7A shown.

[0285] In some possible implementation manners, the computing device according to this application may at least include at least one processor and at least one memory. Among them, the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps in the method for training an information extraction model according to various exemplary implementation manners of this application described above in this specification or the steps in the method for obtaining a knowledge graph. For example, the processor may execute the steps as shown in Figure 3 or Figure 7A shown.

[0286] Next, referring to Figure 12Describe the computing device 120 according to this embodiment of the present application. Figure 12 The computing device 120 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0287] As Figure 12 , the computing device 120 is presented in the form of a general-purpose computing device. The components of the computing device 120 may include, but are not limited to: at least one of the above-mentioned processing units 121, at least one of the above-mentioned storage units 122, and a bus 123 connecting different system components (including the storage unit 122 and the processing unit 121).

[0288] The bus 123 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.

[0289] The storage unit 122 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 1221 and / or a cache storage unit 1222, and may further include a read-only memory (ROM) 1223.

[0290] The storage unit 122 may further include a program / utility 1225 having a set (at least one) of program modules 1224. Such program modules 1224 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0291] The computing device 120 may also communicate with one or more external devices 124 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the computing device 120, and / or communicate with any device that enables the computing device 120 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 125. And, the computing device 120 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 126. As shown in the figure, the network adapter 126 communicates with other modules for the computing device 120 through the bus 123. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 120, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0292] In some possible embodiments, various aspects of the method for training an information extraction model or the method for obtaining a knowledge graph provided in this application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the method for training an information extraction model or the steps in the method for obtaining a knowledge graph according to various exemplary embodiments described above in this specification. For example, the computer device can execute the steps shown in Figure 3 or Figure 7A .

[0293] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0294] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0295] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A method for training an information extraction model, characterized in that, the method comprises: obtaining a training sample set including a non-noise sample set and a noise sample set, wherein each non-noise sample in the non-noise sample set is a sentence pack labeled with a non-noise label, each noise sample in the noise sample set is a sentence pack labeled with a noise label, each sentence pack includes multiple sentences for describing the association information between entities, the sentences are unstructured natural language texts, and the association information includes the relationship between entities or the events involved by entities; training an initial first information extraction model with the noise sample set, and during the training process, updating the first information extraction model based on the difference between the prediction result of the first information extraction model for the noise sample and the noise label corresponding to the noise sample, to obtain a first intermediate model; updating the first information extraction model based on the difference between the prediction result of the first intermediate model for the noise sample and the prediction result of an initial second information extraction model for the non-noise sample, to obtain a second intermediate model, wherein the first information extraction model and the second information extraction model have the same parameters, and the second information extraction model is used to guide the training of the first information extraction model; updating the second intermediate model based on the difference between the prediction result of the second intermediate model for the non-noise sample and the non-noise label corresponding to the non-noise sample, to obtain a trained reference model; adjusting the parameters of the trained reference model based on a preset smoothing coefficient, to obtain a target information extraction model for obtaining entity association information, wherein the entity association information is used to construct a knowledge graph.

2. The method according to claim 1, characterized in that, the noise sample set is obtained by label transfer of the labels of each non-noise sample in the non-noise sample set.

3. The method according to claim 2, characterized in that, the noise sample set includes M, where M is a positive integer; the M noise sample sets obtained by label transfer of the labels of each non-noise sample in the non-noise sample set specifically include: for each non-noise sample in the non-noise sample set, performing M times of label transfer on each non-noise sample according to the non-noise labels of other non-noise samples except itself, to obtain M noise labels corresponding to each non-noise sample; generating M noise samples corresponding to each non-noise sample respectively based on the sentence pack included in each non-noise sample and the M noise labels corresponding to each non-noise sample, the sentence packs of the noise samples corresponding to each non-noise sample are the same, and the noise samples obtained by each non-noise sample through label transfer each time belong to a noise sample set.

4. The method according to claim 3, characterized in that, For each non-noise sample in the non-noise sample set, M times of label transfer are respectively performed on each non-noise sample according to the non-noise labels of other non-noise samples except itself, and M noise labels corresponding to each non-noise sample are obtained, which specifically includes: For any non-noise sample, the following process is executed each time a corresponding noise label is obtained through label transfer: Obtain the similarity between the non-noise sample and other non-noise samples except the non-noise sample, and use the non-noise label of any one of the top N other non-noise samples with the highest similarity ranking as the noise label obtained after label transfer of the non-noise sample, where N is a positive integer.

5. The method according to claim 3, characterized in that, The prediction result of the first information extraction model for the noise sample is the first prediction label; the initial first information extraction model is trained using the noise sample set, and during the training process, the first information extraction model is updated based on the difference between the prediction result of the first information extraction model for the noise sample and the noise label corresponding to the noise sample, and a first intermediate model is obtained, which specifically includes: Input the noise samples in the M noise sample sets into the first information extraction model in batches, and obtain M batches of first prediction labels output by the first information extraction model, where the noise samples in the same batch belong to the same noise sample set, and the noise samples in different batches belong to different noise sample sets; Obtain the first classification loss function determined based on the difference between each batch of first prediction labels and the noise labels of the corresponding noise samples; Perform a gradient update on the first information extraction model respectively according to each first classification loss function to obtain M of the first intermediate models; The first information extraction model is updated based on the difference between the prediction result of the noise sample by the first intermediate model and the prediction result of the non-noise sample by the initial second information extraction model, and a second intermediate model is obtained, which specifically includes: Based on the difference between the prediction result of the noise sample by each first intermediate model and the prediction result of the non-noise sample by the second information extraction model respectively, obtain a fusion difference, and update the first information extraction model according to the fusion difference to obtain the second intermediate model.

6. The method according to claim 5, characterized in that, The prediction result of each first intermediate model for the noise sample is the second prediction label, and the prediction result of the second information extraction model for the non-noise sample is the third prediction label; based on the difference between the prediction result of the noise sample by each first intermediate model and the prediction result of the non-noise sample by the second information extraction model respectively, obtain a fusion difference, and update the first information extraction model according to the fusion difference to obtain the second intermediate model, which specifically includes: Input the noise samples in the M noise sample sets into the corresponding first intermediate state models respectively to obtain the second predicted labels output by each first intermediate state model; and input the non-noise samples in the non-noise sample set into the second information extraction model to obtain the third predicted labels output by the second information extraction model; Determine the consistency loss function based on KL divergence respectively according to the differences between each batch of the second predicted labels and the third predicted labels; Average the M consistency loss functions to obtain the fusion difference; Perform at least one gradient update on the first information extraction model according to the fusion difference to obtain the second intermediate state model, wherein the error between the prediction result of the second intermediate state model and the prediction result of the second information extraction model is within a specified range.

7. The method according to claim 6, wherein, The prediction result of the second intermediate state model for non-noise samples is the fourth predicted label; updating the second intermediate state model based on the difference between the prediction result of the second intermediate state model for non-noise samples and the non-noise label corresponding to the non-noise sample to obtain the trained reference model, specifically including: Input the non-noise samples in the non-noise sample set into the second intermediate state model to obtain the fourth predicted labels output by the second intermediate state model; Obtain the second classification loss function determined based on the difference between the fourth predicted label and the non-noise label of the corresponding non-noise sample; Perform at least one gradient update on the second intermediate state model according to the second classification loss function to obtain the trained reference model, wherein the error between the prediction result of the trained reference model and the prediction result of the second intermediate state model is within a specified range.

8. The method according to any one of claims 1 to 7, wherein, Adjusting the parameters of the trained reference model based on a preset smoothing coefficient to obtain a target information extraction model for obtaining entity association information, specifically including: Performing exponential averaging on the parameters of the trained reference model according to the preset smoothing coefficient to obtain the target information extraction model.

9. A method for obtaining a knowledge graph, wherein, The method includes: Obtain the text to be processed, wherein the text to be processed is an unstructured natural language text for describing the association information between entities, and the association information includes the relationship between entities or the events involved by entities; Input the text to be processed into the trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model, wherein the target information extraction model is trained by the method according to any one of claims 1 to 8; Construct a knowledge graph based on the entity association information.

10. An apparatus for training an information extraction model, wherein, including: An acquisition unit, configured to acquire a training sample set including a non-noise sample set and a noise sample set, where each non-noise sample in the non-noise sample set is a sentence pack with a labeled non-noise label, each noise sample in the noise sample set is a sentence pack with a labeled noise label, each sentence pack includes multiple sentences for describing the association information between entities, the sentences are unstructured natural language texts, and the association information includes the relationship between entities or the events involved by the entities; A first update unit, configured to train an initial first information extraction model using the noise sample set, and during the training process, update the first information extraction model based on the difference between the prediction result of the noise sample by the first information extraction model and the noise label corresponding to the noise sample, to obtain a first intermediate model; A second update unit, configured to update the first information extraction model based on the difference between the prediction result of the noise sample by the first intermediate model and the prediction result of the non-noise sample by an initial second information extraction model, to obtain a second intermediate model, where the first information extraction model and the second information extraction model have the same parameters, and the second information extraction model is used to guide the training of the first information extraction model; A third update unit, configured to update the second intermediate model based on the difference between the prediction result of the non-noise sample by the second intermediate model and the non-noise label corresponding to the non-noise sample, to obtain a trained reference model; An adjustment unit, configured to adjust the parameters of the trained reference model based on a preset smoothing coefficient, to obtain a target information extraction model for obtaining entity association information, where the entity association information is used to construct a knowledge graph.

11. The apparatus according to claim 10, wherein, the noise sample set is obtained by performing label transfer on the labels of each non-noise sample in the non-noise sample set.

12. The apparatus according to claim 10 or 11, wherein, the adjustment unit is specifically configured to: perform exponential averaging on the parameters of the trained reference model according to the preset smoothing coefficient, to obtain the target information extraction model.

13. An apparatus for obtaining a knowledge graph, wherein, comprising: An acquisition unit, configured to acquire a text to be processed, where the text to be processed is an unstructured natural language text for describing the association information between entities, and the association information includes the relationship between entities or the events involved by the entities; An information extraction unit, configured to input the text to be processed into a trained target information extraction model, and extract the entity association information in the text to be processed based on the target information extraction model, where the target information extraction model is obtained by the method according to any one of claims 1 to 8; A construction unit, configured to construct a knowledge graph based on the entity association information.

14. An electronic device, wherein, It includes a processor and a memory. Among them, the memory stores program codes. When the program codes are executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 8 or the steps of claim 9.

15. A computer-readable storage medium, characterized in that it includes program codes. When the program codes run on an electronic device, the program codes are used to cause the electronic device to execute the steps of any one of claims 1 to 8 or the steps of claim 9.

Citation Information

Patent Citations

  • Entity relation extraction method

    CN108733792A