Prompt-based entity-level autoregression relation triple extraction model training method, extraction method and device
By adopting a prompt-based entity-level autoregressive relational triple extraction model in the field of network operation and maintenance, combined with the extended BIO tag mechanism and BERT encoder, the segmented entity problem is solved, and the accuracy and completeness of relational triple extraction is improved.
Patent Information
- Application Number
- CN202510198000.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art is difficult to effectively deal with the segmented entity problems in the field of network operations and maintenance, resulting in the limitation of the accuracy and completeness of relational triple extraction.
A cues-based entity-level autoregressive relationship triple extraction model training method is adopted. Through the combination of the extended BIO tag mechanism and the BERT encoder and classifier, an initial neural network model is built, and the entity pairs and their relationships are predicted one by one during the single-step autoregressive extraction process, and label smoothing technology is used during the training process.
It effectively solves the problem of segmented entities, improves the accuracy and completeness of relational triple extraction, especially in complex scenarios.
Smart Images

Figure CN120216981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of relational triple extraction, and in particular, to a method, an extraction method, and a device for training an entity-level autoregressive relational triple extraction model based on prompts. Background Art
[0002] With the continuous expansion of the scale of communication networks, network operation and maintenance face increasingly complex challenges. Among them, knowledge graph technology is considered to be one of the key means to achieve intelligent network operation and maintenance. As the core technology for constructing knowledge graphs, the main task of relational triple extraction is to extract entities and their relationships from natural language texts to form structured triple data. However, most existing studies focus on the general field, and relatively few studies are aimed at the network operation and maintenance field. Currently, there is a special phenomenon in the corpus of the network operation and maintenance field, that is, the problem of segmented entities, which is relatively rare in the general field. A segmented entity refers to a complete entity composed of multiple discontinuous sequences in the text. This phenomenon is particularly obvious in the professional corpus of the network operation and maintenance field, bringing new challenges to the relational triple extraction technology. Most existing relational triple extraction methods are unable to effectively handle the problem of segmented entities, resulting in limited applications in the professional field. Therefore, how to solve the problem of segmented entities and improve the accuracy and integrity of relational triple extraction has become an urgent technical problem to be solved in the current network operation and maintenance field. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, an extraction method, and a device for training an entity-level autoregressive relational triple extraction model based on prompts to eliminate or improve one or more defects existing in the prior art and solve the problem of identifying segmented entities composed of multiple discontinuous sequences in the relational triple extraction task.
[0004] One aspect of the present invention provides a method for training an entity-level autoregressive relational triple extraction model based on prompts, the method comprising the following steps:
[0005] Obtain a plurality of samples, each sample being a sentence derived from a target topic environment, the sentence being represented as a token sequence; each sample contains a plurality of relational triples, the relational triple including a subject, an object, and the relationship therebetween, the subject and the object being entities composed of one or more entity segments, the entity segment being a subsequence of the sentence containing one or more continuous tokens; among or within the relational triples in one sample, there are three situations: single entity overlap, entity pair overlap, and segmented entity; the single entity overlap means that there is one identical entity between two relational triples, the entity pair overlap means that the entities between two relational triples are the same but the relationships are different, and the segmented entity means that there is an entity within the relational triple composed of multiple non-continuous and non-overlapping entity segments;
[0006] Add labels to each token of the sentences in each sample to construct a training sample set, where the labels distinguish the starting tokens of entity segments, the internal tokens of entity segments, and other non-entity tokens;
[0007] Obtain an initial neural network model, including a BERT encoder and a classifier; in a single-step extraction process, the initial neural network model takes the sentence, as well as the relationship type added at the end of the sentence and the entity pairs that have been extracted, as inputs and outputs the next entity pair. The initial neural network model performs multiple rounds of autoregressive extraction of the prediction results of entity pairs one by one according to the multiple relationship types; construct a cross-entropy loss function based on label smoothing according to the prediction results and the labels, and use the training sample set to update the parameters of the initial neural network model to obtain a target relationship triple extraction model.
[0008] In some embodiments, adding labels to each token of the sentences in each sample to construct a training sample set, where the labels distinguish the starting tokens of entity segments, the internal tokens of entity segments, and other non-entity tokens, includes:
[0009] The starting token of the entity segment includes three parts. The first part marks that the current token belongs to the starting token of the entity segment; the second part marks the entity type, and the entity type includes the subject and the object; the third part marks the segmentation ordinal number of the entity segment represented by the current token within the corresponding entity.
[0010] The internal token of the entity segment includes two parts. The first part marks that the current token belongs to the internal token of the entity segment; the second part marks the entity type, and the entity type includes the subject and the object.
[0011] The other non-entity token includes one part. The first part marks that the current token belongs to the other non-entity token.
[0012] Among them, the first token of the entity segment is marked as the starting token of the entity segment, and all subsequent tokens starting from the second token in the entity segment are marked as the internal tokens of the entity segment. The entity types of the corresponding labels of all tokens in the entity segment are the same.
[0013] In some embodiments, the classifier includes consecutive fully connected layers and a Softmax layer; the BERT encoder is pre-trained using the text data of the target topic environment.
[0014] In some embodiments, the method further includes: pruning the attention heads and layers of the BERT encoder; introducing a local attention mechanism or a cross-segment attention mechanism into the BERT encoder.
[0015] In some embodiments, the BERT encoder adopts DistilBERT / TinyBERT.
[0016] In some embodiments, a cross-entropy loss function is constructed based on the prediction result and the label on the basis of performing label smoothing, and the expression of the loss is:
[0017]
[0018] where y LS represents the smoothed label, where j represents the index of each class in the vector y LS and j0 represents the index of the actual class, and α is a hyperparameter representing the degree of smoothing.
[0019] In some embodiments, the method further includes: dividing the training sample set into a training set, a validation set, and a test set, using the training set to perform model training, using the validation set to evaluate and optimize the target relation triple extraction model, using the test set to calculate the standard accuracy, recall rate, and F1 score of the trained and optimized target relation triple extraction model, and generating a performance evaluation report.
[0020] On the other hand, the present invention also provides a prompt-based entity-level autoregressive relation triple extraction method, and the method includes the following steps:
[0021] Obtain the sentence to be recognized related to the target topic environment;
[0022] Tokenize the sentence to be recognized and input it into the target relation triple extraction model in the above-mentioned prompt-based entity-level autoregressive relation triple extraction model training method, and output the entity recognition result.
[0023] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0024] On the other hand, the present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0025] The beneficial effects of the present invention are at least:
[0026] The training method, extraction method and device of the prompt-based entity-level autoregressive relation triple extraction model according to the present invention, in the complex relation scenarios with single entity overlap, entity pair overlap and segmented entities, through an extended BIO tag mechanism, add multiple tags to multiple entity segments of sentences in each sample according to multiple entity types to construct a training sample set, and the tags distinguish the starting tokens, internal tokens and other non-entity tokens of sentences in each sample. An initial neural network model is constructed using a BERT encoder and a classifier. In the single-step autoregressive extraction process, the model takes the sentence, relation type and the extracted entity pair as inputs, and predicts entity pairs and their relations one by one through multiple rounds of autoregression for different relation types. In the training process, label smoothing technology is adopted, and the parameters are updated in combination with the cross-entropy loss function to optimize the extraction performance of the model. The present invention can efficiently label segmented entities by means of multi-segmentation, multi-relation categories and distinguishing intra-sentence attributes. In the model learning process, entities are predicted based on multiple rounds of autoregression, which can effectively solve the problem of relation triple extraction in complex situations with segmented entities.
[0027] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and in part will become obvious to those skilled in the art upon examination of the following or may be learned from the practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structure particularly pointed out in the specification as well as the drawings.
[0028] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. Brief Description of the Drawings
[0029] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not limit the present invention. In the drawings:
[0030] Figure 1 It is a schematic diagram of the multi-step autoregressive extraction process of the model in the training method of the prompt-based entity-level autoregressive relation triple extraction model according to an embodiment of the present invention.
[0031] Figure 2 It is a schematic diagram of the single-step extraction process of the model in the training method of the prompt-based entity-level autoregressive relation triple extraction model according to an embodiment of the present invention.
[0032] Figure 3 It is an example diagram of a triple that is judged to be correct although ROUGE-1 < 1.
[0033] Figure 4The distribution diagram of the ROUGE-1 scores for multiple incompletely matching predicted triples.
[0034] Figure 5(a) shows the F1 scores of PEAR and other LLM-based baselines under different numbers of triples.
[0035] Figure 5(b) shows the F1 scores of PEAR and other LLM-based baselines under different numbers of Lingpai entities.
[0036] Figure 5(c) shows the recall rates of PEAR and other LLM-based baselines under different entity score segments. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0038] Herein, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0039] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0040] Among numerous enabling technologies, the knowledge graph, with its unique advantages, has already become one of the key strategies for realizing intelligent network operation and maintenance, attracting extensive attention in the academic community. It should be emphasized that the quality of the knowledge graph directly determines its actual effectiveness in downstream network operation and maintenance tasks. Currently, most knowledge graphs used in network operation and maintenance-related research mainly originate from the construction of structured data, which to a certain extent limits the diversity of data sources and lacks unstructured data such as reasoning knowledge. To achieve a higher level of intelligent goals, it is particularly crucial to mine high-quality knowledge from unstructured texts.
[0041] The relationship triple extraction (RTE) technology is dedicated to simultaneously identifying entity pairs and their associated relationships from unstructured texts and is one of the core technologies for constructing knowledge graphs. The triples extracted by this technology are presented in the format of "<subject, relationship, object>" and constitute the basic units of the knowledge graph. With the gradual in-depth research on relationship triple extraction, its development context can be roughly divided into two major camps: non-LLM (large language model) methods and LLM-based methods.
[0042] Non-LLM methods usually encode text data with an encoder and then use methods such as sequence labeling, table filling, or decoder generation to obtain information related to triples. Against the backdrop of the rapid development of LLMs, the method of relation triple extraction based on LLMs generally adds prompt information before and after the text, inputs it into the LLM, and finally guides the LLM to output a formatted triple result. These two methods have their own advantages and also have certain limitations. In the general domain corpus, the current state-of-the-art non-LLM methods often outperform the LLM-based methods. However, the LLM-based methods, with their vast pre-trained knowledge reserves and the flexibility of the output content, show stronger potential when dealing with complex and unique scenarios.
[0043] Although there have been many exciting results in the field of relation triple extraction, unfortunately, most of them focus on extracting triples from traditional complex scenarios, such as single entity overlap (SEO) and entity pair overlap (EPO) scenarios.
[0044] However, in the Chinese network operation and maintenance corpus, there is an easily overlooked situation: a complete entity may be composed of multiple discontinuous segments in the text. For example, in the sentence "The means of controlling cell coverage include adjusting the MS received signal level and the RACH access level.", the entity "adjust the RACH access level" needs to combine two discontinuous text segments "adjust" and "RACH access level". In this application, this scenario is called a "segmented entity". The reason for ignoring segmented entities is mainly that most studies focus on the general domain, where entities can usually be represented by concise phrases such as "person name", "location", etc. In addition, widely used datasets, such as NYT, WebNLG, etc., rarely involve the situation of segmented entities.
[0045] In contrast, in the corpus of professional fields such as network operation and maintenance, entities are often longer and more complex, making the phenomenon of segmented entities more prominent. Network operation and maintenance refers to a series of management activities carried out to ensure the safe and efficient operation of telecommunication network services, and its core tasks cover multiple aspects such as equipment management, network monitoring, and fault analysis. In recent years, the rapid expansion of the communication network scale has brought many challenges to network operation and maintenance, mainly manifested as resource-intensive repetitive tasks, a high degree of dependence on professional knowledge, and fragmented knowledge acquisition. Currently, most existing non-LLM methods lack an effective mechanism for dealing with segmented entities. Although the LLM-based methods have recently shown the potential to solve this problem by generating complete entity texts, the continued use of early datasets has limited the researchers' full understanding and prioritized handling of this problem.
[0046] In view of this, the present invention provides a training method for a prompt-based entity-level autoregressive relation triple extraction model, and the method includes the following steps S101 to S103:
[0047] Step S101: Obtain a plurality of samples, each sample being a sentence originating from a target topic environment, and the sentence being represented as a token sequence; each sample contains a plurality of relation triples, and a relation triple includes a subject, an object, and the relation therebetween. The subject and the object are entities composed of one or more entity segments, and an entity segment is a subsequence of the sentence containing one or more consecutive tokens; there are three cases of single entity overlap, entity pair overlap, and segmented entity among or within the relation triples in one sample; single entity overlap means that there is one same entity between two relation triples, entity pair overlap means that the entities between two relation triples are the same but the relations are different, and segmented entity means that there is an entity within a relation triple composed of multiple non-consecutive and non-overlapping entity segments.
[0048] Step S102: Add labels to each token of the sentences in each sample to construct a training sample set, and the labels distinguish the starting token of the entity segment, the internal token of the entity segment, and other non-entity tokens.
[0049] Step S103: Obtain an initial neural network model, including a BERT encoder and a classifier; in the single-step extraction process of the initial neural network model, the sentence, the relation type added at the end of the sentence, and the extracted entity pair are used as inputs and the next entity pair is output. The initial neural network model performs multiple rounds of autoregression according to multiple relation types to extract the prediction results of entity pairs one by one; based on the prediction results and the labels, a cross-entropy loss function is constructed on the basis of performing label smoothing, and the initial neural network model is updated with the training sample set to obtain the target relation triple extraction model.
[0050] In step S101, for the samples to be obtained, it can be targeted at a specific language environment, so a target topic environment can be set to obtain corresponding samples. Exemplarily, it can be collected for the network operation and maintenance environment. When configuring the samples, various cases of relation triples are arranged to ensure the learning quality. In this embodiment, the relation triple types in the samples include at least three types: single entity overlap, entity pair overlap, and / or segmented entity.
[0051] The goal of the relation triple extraction task is to extract all possible triples from a given sentence S = {w1, w2,..., w L} with L tokens (tokens) and combine them into format. Among them, r i represents a relation type selected from a predefined set, s i and o iDenote the subject and the object respectively, both of which are composed of sequences of w j ∈ S, and there is an r i relationship between them.
[0052] The complex scenarios of the relation triple extraction task mainly include the following:
[0053] Single entity overlap (SE0), that is, there are two triples (s1, r1, o1) and (s2, r2, o2), where s1 or o1 is equal to s2 or o2.
[0054] Entity pair overlap (EP0), that is, there are two triples (s1, r1, o1) and (s2, r2, o2), where the set {s1, o1} is equal to {s2, o2}, but r1 is not equal to r2.
[0055] Segmented entity, that is, there is an entity that can be represented, where n1 > 0, n2 > 0, j ≠ i + n1 + 1, and the intervals [i, i + n1] and [j, j + n2] do not overlap.
[0056] In step S102, labels are added to the samples to prepare for subsequent model training. In some embodiments, labels are added to each token of the sentences in each sample to construct a training sample set. The labels distinguish the start token of the entity segment, the internal token of the entity segment, and other non-entity tokens, including:
[0057] The start token of the entity segment includes three parts. The first part marks that the current token belongs to the start token of the entity segment; the second part marks the entity type, and the entity types include the subject and the object; the third part marks the segmented ordinal number of the entity segment represented by the current token within the corresponding entity.
[0058] The internal token of the entity segment includes two parts. The first part marks that the current token belongs to the internal token of the entity segment; the second part marks the entity type, and the entity types include the subject and the object.
[0059] Other non-entity tokens include one part. The first part marks that the current token belongs to other non-entity tokens;
[0060] Among them, the first token of the entity segment is marked as the start token of the entity segment, and all subsequent tokens starting from the second token in the entity segment are marked as internal tokens of the entity segment. The entity types of all tokens corresponding to the labels in the entity segment are the same.
[0061] In order to solve the problem of segmented entities, this application proposes an improved tagging mechanism - segmented BIO. In this mechanism, each entity type is assigned at most (m+1) tags - "B1, B2, ..., I", where each B tag represents the first token of each entity segment. If the actual number of segments n of the entity is less than m, only "B1, B2, ..., Bn, I" will be used. In the process of converting from labels to entity text, the entity segments are connected in ascending order according to the subscript numbers of their "B" tags. Finally, combined with the above description, the tag set of the segmented BIO tagging mechanism is given:
[0062] T set ={SUBJ-B1,SUBJ-B2,…,SUBJ-B m ,SUBJ-I,
[0063] OBJ-B1,OBJ-B2,…,OBJ-B m ,OBJ-I,O}
[0064] Among them, T set The length is equal to d t .
[0065] The present invention uses segmented, multi-labeled forms to add tags and distinguish tokens according to the attributes within the sentence.
[0066] Introducing multi-segment start markers (such as SUBJ-B1, SUBJ-B2) to clearly identify the starting positions of different segments, provide richer structural information, and allow the model to directly identify discontinuous entity segments through labels and combine them into complete entities. In scenarios involving segmented entities (such as network operation and maintenance), the model recall rate is significantly improved. Experiments show that segmented BIO contributes to a 7.1% improvement in the F1 score. When irrelevant content is interspersed between entity segments (such as "adjusting the MS receive signal level and RACH access level"), segmented BIO can clearly distinguish which tokens belong to different segments of the same entity to avoid misjudgment. It can adapt to diverse scenarios and flexibly support different numbers of segments (such as 2-segment and 3-segment entities), and cover more complex entity structures through label design. By clearly marking the boundaries of entity segments, the model can still accurately extract the target entity in noisy text (such as redundant descriptions and nested entities), improving anti-interference capabilities.
[0067] In step S103, in order to achieve the goal of the relation triple extraction task and overcome the complex situations involving single entity overlap, entity pair overlap and / or segmented entities, the present invention provides a prompt-based entity-level autoregressive relation triple extraction method, hereinafter referred to as PEAR.
[0068] Figure 1 and Figure 2 The PEAR framework is shown, Figure 1Illustrates the extraction process of PEAR, which involves iteratively extracting entity pairs from a given sentence. Figure 2 Illustrates the single-step extraction process in PEAR and the structure of the model in PEAR. The single-step extraction process is formulated as a sequence tagging task. Specifically, a sentence with prompts is mapped to a sequence and then input into the model, which then assigns labels to each token in the original sentence. Figure 2 The output label sequence in is a typical example of segment BIO. This mechanism can accurately label segmented entities, enabling the model to extract these entities. In addition, a label smoothing mechanism is introduced in the training phase to prevent the model from overfitting.
[0069] The framework of PEAR can be represented by the following formula:
[0070] (s k+1 , o k+1 ) = Model(S, r, (s k , o k ))
[0071] k = 0, 1, 2, …, N r
[0072]
[0073] where "Model" represents the single-step extraction process in PEAR. "Model" takes as input "sentence S", "relation type r", and "the k-th entity pair (s k , o k )" related to r, and then outputs (s k+1 , o k+1 ). This extraction process will be iteratively executed to form a "entity-level autoregressive" structure. In the initial state, (s0, o0) is empty, indicating the extraction of the first entity pair; when another empty entity pair is obtained, the process terminates. It should be noted that the total number N of entity pairs is unknown before r is obtained. In addition, a single-round autoregressive process can only extract entity pairs belonging to a specific relation type. To comprehensively extract all relation triples, PEAR needs to access all relation types (assuming a total of M), and execute M different autoregressive processes. By fusing input information through a prompt-based method, specifically, two text descriptions are appended to the end of the original sentence, representing "relation type" and "previously extracted entity pairs" respectively, so that the sentence and the prompt can be input into the model in a unified format.
[0074] The present invention uses BERT to encode the input tokens and generate a sequence of feature vectors. The BERT encoder can be pre-trained using the text data of the target topic environment. First, the sentence S and its prompt S p ={w L+1 , w L+2 , …, w L′} are tokenized, and then these tokens are mapped into a low-dimensional space to generate a sequence of word embeddings X. Then this sequence is input into BERT, and BERT subsequently outputs a sequence of feature vectors H. The process in BERT can be simply described by the following formula:
[0075] H = BERT(X; Θ1)
[0076] X = {x1, x2, …, x L′}
[0077] H = {h1, h2, …, h L′}
[0078] where the dimensions of both X and H are L′×d. L′ represents the number of tokens in the sentence after adding the prompt, and L′>L. d represents the dimension of the BERT hidden layer. The symbol Θ1 represents all the trainable parameters in BERT. In some embodiments, pruning can also be performed on the attention heads and layers of the BERT encoder; a local attention mechanism or a cross-segment attention mechanism can be introduced into the BERT encoder. To achieve lightweight, the BERT encoder can also adopt DistilBERT / TinyBERT.
[0079] Subsequently, a fully connected layer and Softmax are used to obtain the sequence T to represent the probabilities of the tokens, which can be described by the following formula:
[0080] T = {t1, t2, …, t L}
[0081] t i = Softmax((h i ·W + b)
[0082] where and represent the trainable parameters of the fully connected layer. The vector t i corresponds to the probabilities of each token and has a dimension of d t . It should be noted that the range of the index i is from 1 to L, corresponding to the original sentence part.
[0083] During the training process, the cross-entropy loss function is used to calculate the loss between the predicted labels and the true labels. In addition, a label smoothing mechanism is introduced to prevent the model from overfitting, especially when the training samples are limited. This move aims to improve the robustness of PEAR.
[0084] In the formula of cross-entropy loss, the label y is usually represented as a one-hot vector. Label smoothing is a regularization technique that prevents the model from being too confident in its predictions during training by reducing the values that were originally 1 in y and increasing the values that were originally 0 in y. The label y after label smoothing LS can be expressed as:
[0085]
[0086]
[0087] where j represents the index of each class in the vector y LS and j0 represents the index of the actual class, and α is a hyperparameter representing the degree of smoothing.
[0088] The loss function after label smoothing can be expressed as:
[0089]
[0090] In some embodiments, the method further includes: dividing the training sample set into a training set, a validation set, and a test set, performing model training using the training set, evaluating and tuning the target relation triple extraction model using the validation set, calculating the standard accuracy, recall rate, and F1 score for the target relation triple extraction model after training and tuning using the test set, and generating a performance evaluation report.
[0091] On the other hand, the present invention also provides a prompt-based entity-level autoregressive relation triple extraction method, and the method includes the following steps S201 to S202:
[0092] Step S201: Obtain the sentence to be recognized related to the target topic environment.
[0093] Step S202: Tokenize the sentence to be recognized and input it into the target relation triple extraction model in the above-mentioned prompt-based entity-level autoregressive relation triple extraction model training method, and output the entity recognition result.
[0094] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0095] On the other hand, the present invention also provides a computer program product, including computer programs / instructions, which implement the steps of the above method when executed by a processor.
[0096] The effects of the present invention will be described below in conjunction with specific embodiments:
[0097] 1. Dataset construction process
[0098] In this embodiment, the experiment is carried out on the constructed CMIM23-NOM1-RA (China Mobile Innovator Marathon 2023 - Network Operation and Maintenance, task 1 - Re Annotation) dataset. The original text of this dataset comes from task 1 of the "2023 China Mobile Maker Marathon - Network Operation and Maintenance Q&A Based on Knowledge Graph" organized by China Mobile in 2023, and is accompanied by re-annotations based on the structure of the present invention. The purpose of task 1 of this competition is to extract relationship triples from the Chinese network operation and maintenance corpus for constructing a domain knowledge graph to complete subsequent tasks. The original dataset of task 1 can be called "CMIM23-NOM1", which contains a training set and a test set. The training set contains 4354 labeled samples, plus 21206 unlabeled samples, and contestants can freely use them for various purposes. On the other hand, the test set includes 981 unlabeled samples for evaluating the effectiveness of the methods proposed by contestants.
[0099] However, the annotation of CMIM23-NOM1 follows the 0IE (Open Information Extraction task) format. 0IE must extract all structured information including relationship types from the corpus, rather than pre-defining a set of relationship types like in the case of limited-domain relationship triple extraction. Therefore, the labeled training data in CMIM23-NOM1 contains a large number of different relationship types, reaching a total of 547. OIE has rich relationship diversity, can discover new relationship types, and shows good transferability. It is undeniable that the establishment of CMIM23-NOM1 can promote the effective transfer of open-domain models from the general domain to the network operation and maintenance domain. However, open-domain research is inherently troubled by challenges such as difficult annotation, limited accuracy, and the need for a large amount of labeled data. Just having more than 4354 annotated samples cannot fully meet the demand for a large amount of annotation in open-domain research, so the constructed knowledge graph may lack accuracy.
[0100] Therefore, in this embodiment, a dataset for extracting qualified domain relation triples is constructed. A labeling team consisting of five members from information and communication engineering or related disciplines was established. Then, 4354 sentences that had been labeled in the training set of CMIM23-NOM1 and 981 sentences in the test set were relabeled. Each sample was independently labeled by two members and then verified and determined by a third member. During this process, the relation types were adjusted multiple times, and some abnormal samples were processed. Finally, 13 different relation types common in the network operation and maintenance corpus were summarized, and 5337 relabeled samples were obtained, which contained 19509 triples. Finally, 3337 of these samples were set aside for training, 1000 for validation, and 1000 for testing, and it was called the "CMIM23-NOM1-RA" dataset.
[0101] CMIM23-NOM1-RA has the following unique features: The relation types in the dataset exhibit obvious properties of the professional domain. The relations existing in the corpus are divided into 13 types, effectively covering most of the common logical relations in the network operation and maintenance corpus. This dataset is information-intensive, with an average of 3.7 triples per sample. Approximately 17% of the sentences contain more than 5 triples, and in the most extreme case, there are 50 triples. It is worth noting that approximately 18% of the triples involve segmented entities, and the number of segments mainly concentrates in the range of 2 to 4. These features highlight its complexity and pose significant challenges to the relation triple extraction method. Finally, the specific relation types and additional features of the CMIM23-NOM1-RA dataset are shown in Table 1.
[0102] Table 1 Relation types and additional features of the CMIM23-NOM1-RA dataset
[0103]
[0104] 2. Experimental settings
[0105] Six non-LLM methods and three LLM-based methods were selected as baselines. The following content will describe the implementation of these baselines and the evaluation metrics used in the experiment.
[0106] 2.1 Implementation of non-LLM baselines
[0107] Six non-LLM baselines were replicated, namely TPLinker, SPN, PRGC, OneRel, BiRTE, and UniREL. They are all state-of-the-art BERT-based methods in recent years. To adapt them to CMIM23-NOM1-RA, the following adjustments were made in this embodiment:
[0108] Pretrained model parameters: To ensure compatibility with the Chinese corpus, the pre-trained parameters of these baselines were adjusted to "Chinese-BERT-WWM-EXT" to achieve better Chinese understanding and align them with the parameters used in PEAR.
[0109] Segmented entity processing: Given that these baselines cannot perform the extraction of segmented entities, triples containing segmented entities were excluded from the training set, and the remaining data was used to train their models.
[0110] 2.2 Implementation of LLM-based baselines
[0111] This embodiment also conducted experiments using three high-performance LLMs, namely T5, GLM, and Qwen. To achieve this goal, the following strategies were implemented with reference to Table 2:
[0112] Pretrained model parameters: For T5, after experimental comparison, "Randeng-T5-784M-MultiTask-Chinese" was selected as the pre-trained parameter, which was abbreviated as "Rd-T5-784MTC" in the experiment. For GLM and Qwen, due to hardware limitations, "GLM4-9B-Chat" and "Qwen2-7B" were selected respectively. Although they have fewer parameters, they have excellent performance.
[0113] Data format: CMIM23-NOM1-RA was adjusted to the question-and-answer form to adapt to the paradigm of LLMs, and appropriate prompt information was added. For an example of the question-and-answer format, please refer to Appendix a.
[0114] Fine-tuning: Two fine-tuning strategies were adopted. For T5, all parameters were fine-tuned to achieve the best performance. For GLM and Qwen, LoRA fine-tuning was used, which enabled satisfactory results to be obtained under limited resources.
[0115] Data augmentation: Data augmentation operations were carried out. Specifically, each sample in the dataset was expanded into 10 question-and-answer pairs, providing different examples for each question. During the validation process, the triples predicted by more than half of the expanded samples were selected as the final predictions of this embodiment. By adopting this operation, an improvement in performance was found in this embodiment.
[0116] In addition, the implementation details of PEAR and all baselines are presented in Table 2. All LLM-based baselines and the experiments of OneRe1 were conducted on NVIDIA A800 80GB GPUs, and other experiments including PEAR were conducted on NVIDIA V100 32GB GPUs.
[0117] Table 2 Implementation details of PEAR and all baselines.
[0118] In the table, "*" indicates that the setting of this item is the same as that in the baseline paper, and "-" indicates that this item does not need to be set.
[0119]
[0120] 3. Evaluation Metrics
[0121] Two triple judgment criteria are adopted to judge the correctness of triples. One is looser, namely "ROUGE matching", and the other is more strict, namely "exact matching".
[0122] The first criterion is different from previous work. In previous studies, the correctness of entities depended entirely on matching the first or last token, which was a loose criterion. However, in this paper, the entities are longer and the scenarios of segmented entities are complex, making the traditional criterion inapplicable. Therefore, this embodiment proposes a loose criterion that includes ROUGE. ROUGE is a widely used metric in text summarization to measure the similarity between the text generated by the model and the reference text, and the higher the value, the higher the similarity. In this experiment, "ROUGE matching" is defined. If the ROUGE-1 score between the predicted entity and the ground truth is greater than 0.6, the entity is considered correct. If the relation type, as well as the subject and object, are all correct, the triple is considered correct. ROUGE matching is used to tolerate entities with only a small number of unmatched tokens but the same basic meaning, and entities that do not hinder the construction of the knowledge graph. In addition, this judgment criterion with a certain degree of tolerance can demonstrate the effectiveness of the relation triple extraction method in scenarios closer to the real world.
[0123] The second criterion - "exact matching" is more strict because it requires that the predicted entity must be the same as the ground truth to be considered correct. This criterion reflects the exact extraction ability of the relation triple extraction method.
[0124] After determining the above two criteria, the standard accuracy (Prec.), recall (Rec.), and F1 score are reported respectively to comprehensively evaluate the performance of the method.
[0125] 4. Experimental Results
[0126] For the experiment configured based on the above method, the results are shown in Table 3.
[0127] Table 3 Experimental Results Table
[0128]
[0129] Table 3 shows the "exact match", "ROUGE match", and "segmented entity" results of PEAR and all other baselines on the CMIM23-NOM1-RA dataset. "Exact match" and "ROUGE match" provide statistics based on all tokens and predicted triples. "Segmented entity" represents the exact match results when only triples containing segmented entities are retained.
[0130] The "exact match" and "rough match" columns in Table 3 show the comprehensive performance of the methods. It is obvious that PEAR outperforms other baselines in most metrics. Under the condition of exact match, the F1 score of PEAR is 0.4% higher than that of the best-performing UniRel among the baselines and 1.3% higher than that of the best method GLM4-9B-Chat based on LLM. Under the condition of ROUGE match, the F1 score of PEAR is also 2.8% higher than that of the best baseline Rd-T5-784MTC. The significant advantage of PEAR under ROUGE match is mainly attributed to its high recall rate, which greatly improves the F1 score. These results clearly show that in the Chinese network operation and maintenance corpus, PEAR performs better than other general domain methods. Due to the high challenge of the dataset, the scores under "exact match" are relatively low.
[0131] However, it is worth noting that in ROUGE match, the F1 score of PEAR exceeds 70%. This score is a remarkable achievement, especially in a professional field with limited corpus sources and a small training set size. In addition, ROUGE match better reflects the performance of PEAR in real-world applications, and a score of 70.9% can meet the requirements for building a high-quality knowledge graph.
[0132] The "segmented entity" column in Table 3 shows the comparison between PEAR and baselines that can extract segmented entities. As a novel and challenging problem, the scores reported under the condition of segmented entities are not high. The reason is that the extraction of segmented entities not only requires extracting multiple segments from the text but also integrating the text, which poses a major challenge to the model's text understanding ability. In addition, the limited number of training samples with segmented entities also leads to low scores. However, in comparison, PEAR still outperforms all baselines in terms of precision and F1 score.
[0133] 5. Experimental Analysis
[0134] 5.1 Reliability Analysis Based on ROUGE Match
[0135] The ROUGE-1 score is used for analysis. The reliability of "ROUGE match" is explained below, and a statistical-based explanation for setting the metric threshold to 0.6 is provided. ROUGE aims to measure the similarity between two texts, and the higher the ROUGE score, the greater the similarity. Figure 3Two typical triple cases are shown. Although they do not match exactly, they are considered correct. In example (1), compared with the label, the predicted subject contains additional information, but this does not change the meaning of the entity, so it is considered correct. This situation of adding or missing parts is quite common. In example (2), the predicted object of the label truth value is split into two separate parts, resulting in two output triples. In the context of constructing a knowledge graph, this does not affect the integrity of the graph, so it is also considered correct. For those incorrect triples, they often have only a few tokens overlapping with the label, resulting in a significant change in the entity meaning. These triples tend to have a lower ROUGE score.
[0136] To gain a deeper understanding of the ROUGE-1 score distribution of correct and incorrect triples, nearly 200 predicted triples with imperfect matches were further evaluated. The results are as Figure 4 shown. It can be observed that most of the correct triples have scores above 0.7, while most of the incorrect triples have scores below 0.5. Assuming that these two distributions are independent, we set the threshold to 0.6 according to naive Bayes classification. The above cases and statistics indicate that "ROUGE matching" is reliable and the threshold setting is reasonable.
[0137] 5.2. PEAR Advantage Analysis
[0138] Statistics were conducted on the extraction of samples containing different numbers of triples and samples involving SEO and EPO scenarios. The results are shown in Table 4.
[0139] Table 4 F1 scores of each method in samples with different numbers of triples and overlapping patterns
[0140]
[0141]
[0142] The samples were divided into six groups according to the number of triples contained in the samples, and the ranges of the number of triples were from 1 to 5 and ≥6 respectively. The more triples there are in the sample, the higher the complexity of the sentence. The results show that for samples with 3 or fewer triples, PEAR obtains lower F1 scores than OneRel and UniRel. However, when the number of triples is greater than or equal to 4, PEAR obtains the highest scores. At the same time, the F1 scores of each method in ordinary, SEO, and EPO scenarios were also statistically analyzed. It can be observed that PEAR obtains the highest scores in the complex scenarios of SEO and EPO, while UniRel scores the highest in the ordinary scenario.
[0143] The excellent performance of OneRel and UniRel in simple scenarios may stem from their better optimization of accuracy. In the relatively simple scenarios mentioned above, the recall rates among different methods are very close, and OneRel and UniRel with higher accuracy have greater advantages.
[0144] However, PEAR has demonstrated excellent performance in extracting information from more complex texts. We summarize two reasons. First, the combination of the autoregressive mechanism and the segmented BIO mechanism makes PEAR more suitable for extracting entity information, including segmented entities, in complex scenarios. Second, in PEAR, the extraction of different relation types is completely independent, making it unaffected by the presence of multiple relations in a sentence.
[0145] Since complex texts often carry more information, compared with all other baselines, the advantage of PEAR in processing complex texts enables it to obtain the highest F1 score in Table 3.
[0146] 5.3 Analysis of the Effect of Extracting Segmented Entities
[0147] A detailed analysis was conducted on the performance of extracting segmented entities of PEAR and other LLM-based baselines. The aim was to understand the specific factors limiting their scores. This was analyzed from three perspectives: the range of the number of triples in the samples, the range of the number of tokens of entities, and the range of the number of segments within segmented entities.
[0148] First, as can be seen from Figure 5(a), the extraction performance of all methods generally decreases as the number of triples increases. This indicates that the complexity of the text is one of the factors limiting the ability of the method to extract segmented entities.
[0149] Second, Figure 5(b) shows that as the number of entity tokens increases, the performance of the three LLM-based methods shows a slight downward trend, while PEAR is not affected. This is because the generation mechanism of LLM is affected by the output length, while PEAR does not have this problem. In short, the number of entity tokens also has a certain impact on the extraction of segmented entities.
[0150] Finally, in Figure 5(c), it is obvious that the correctly extracted triples are mainly concentrated in the case where the maximum number of segments is 2, and the recall rate of entity extraction divided into more than 3 segments is relatively low. Note that since the total number of triples with 4 segments is only 4, the 25% recall rate of Qwen2-7B means that one triple with a 4-segment entity is correctly extracted. We believe that the increase in the number of segments makes the extraction difficulty increase suddenly, and the number of training samples decreases suddenly, and these two factors jointly limit the current performance of extracting segmented entities.
[0151] 5.4 Ablation Study
[0152] An ablation study was conducted to investigate the contributions of some modules to PEAR. Specifically, ablation experiments were performed on the relation information fusion method and label smoothing mechanism mentioned in Section 2. The results are listed in the last two rows of Table 3.
[0153] PEAR spe_tok_rela means that custom special tokens are assigned to various relation types in the input fusion stage instead of using natural language text to represent them. The results show that representing relations in natural language form can improve the performance of PEAR in exact match, ROUGE match, and segmented entities by 1.8%, 1.5%, and 2.8% respectively. This fully verifies the importance of leveraging the semantic information of pre-trained models.
[0154] PEAR no_seg_ent indicates that PEAR does not contain "segmented BIO", so it cannot extract segmented entities. During the training process, we processed the dataset in the same way as when training non-LLM baselines, that is, all triples containing segmented entities were removed. The results show that the accuracy of PEAR no_seg_ent increased, while the recall decreased, and the overall F1 score was lower than that of PEAR (1.5% lower in exact match and 7.1% lower in ROUGE match). In addition, the performance of PEAR no_seg_ent was closer to that of non-LLM baselines and failed to exceed some excellent baselines. The above phenomena indicate that the "segmented BIO" mechanism significantly improves the performance of PEAR, especially its recall ability.
[0155] PEAR no_LS indicates that the label smoothing mechanism was not used in the loss function. The experimental results show that the method using the cross-entropy loss function with label smoothing is 1.7% higher in exact match and 0.5% higher in ROUGE match and segmented entity cases than using only the conventional cross-entropy loss. This shows that the label smoothing mechanism can alleviate the overfitting phenomenon and has a certain optimization effect on model training.
[0156] To achieve intelligent network operation and maintenance and realize the automated construction of high-quality knowledge graphs in this field, the present invention focuses on researching the relational triple extraction technology for Chinese network operation and maintenance corpora and determines a new complex scenario - segmented entities. To solve the problems brought by segmented entities, the present invention proposes a prompt-based entity-level autoregressive relational triple extraction method called "PEAR". This method is an autoregressive framework that includes an end-to-end sequence tagging model for extracting relational triples. In the tagging stage, the "segmented BI0" tagging scheme is introduced to effectively handle segmented entities. In addition, label smoothing is introduced in the training stage to mitigate overfitting. To further evaluate the performance of PEAR, CMIM23-NOM1-RA is constructed, which is the first high-quality domain-specific relational triple extraction dataset in the network operation and maintenance field. This dataset has the characteristics of strong professionalism in relation types, high information density, and complex scenarios including segmented entities, presenting significant challenges. Finally, extensive experiments and comprehensive analyses are carried out, demonstrating that PEAR achieves the best performance in Chinese network operation and maintenance corpora.
[0157] In summary, for the training method, extraction method, and device of the prompt-based entity-level autoregressive relational triple extraction model described in the present invention, in complex relational scenarios with single-entity overlap, entity-pair overlap, and segmented entities, through the extended BIO tagging mechanism, multiple labels are added to multiple entity segments of the sentences in each sample according to multiple entity types to construct a training sample set, and the labels distinguish the start tokens, internal tokens, and other non-entity tokens of the sentences in each sample. An initial neural network model is constructed using a BERT encoder and a classifier. In the single-step autoregressive extraction process, the model takes the sentence, relation type, and the extracted entity pair as inputs, and for different relation types, predicts the entity pair and its relation one by one through multiple rounds of autoregression. In the training process, the label smoothing technique is adopted, and the parameters are updated in combination with the cross-entropy loss function to optimize the extraction performance of the model. By means of multi-segmentation, multi-relation categories, and distinguishing in-domain attributes, the present invention can efficiently label segmented entities. During the model learning process, based on multiple rounds of autoregression to predict entities, it can effectively solve the problem of relational triple extraction in complex situations with segmented entities.
[0158] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link.
[0159] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0160] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0161] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A prompt-based entity-level autoregressive relation triple extraction model training method, characterized in that: The method comprises the following steps: Acquire multiple samples, each sample is a sentence originating from a target subject environment, and the sentence is represented as a token sequence; each sample contains multiple relationship triples, and the relationship triples contain a subject, an object, and the relationship between them, and the subject and the object are entities composed of one or more entity segments, and the entity segment is a subsequence containing one or more continuous tokens in the sentence; there are three situations between or within the relationship triples in a sample: single entity overlap, entity pair overlap, and segmented entity; the single entity overlap indicates that there is an identical entity between the two relationship triples, the entity pair overlap indicates that the entities between the two relationship triples are identical but the relationship is different, and the segmented entity indicates that there is an entity composed of multiple mutually discontinuous and non-overlapping entity segments in the relationship triple; A training sample set is constructed by adding a label to each token in the sentence of each sample, wherein the label distinguishes the entity segment start token, the entity segment internal token, and other non-entity tokens; An initial neural network model is obtained, including a BERT encoder and a classifier; in a single-step extraction process, the initial neural network model takes the sentence, the relationship type added at the end of the sentence, and the extracted entity pair as input and outputs the next entity pair, and the initial neural network model implements multiple rounds of autoregression according to multiple relationship types to extract prediction results of entity pairs one by one; a cross entropy loss function is constructed based on the prediction results and the labels and label smoothing is performed, and the parameters of the initial neural network model are updated using the training sample set to obtain a target relationship triple extraction model.
2. The prompt-based entity-level autoregressive relationship triple extraction model training method according to claim 1, characterized in that: A training sample set is constructed by adding a label to each token in each sentence of each sample. The label distinguishes entity segment start tokens, entity segment internal tokens, and other non-entity tokens, including: The entity segment start token includes three parts. The first part marks that the current token belongs to the entity segment start token; the second part marks the entity type, which includes subject and object; the third part marks the segment ordinal number of the entity segment represented by the current token in the corresponding entity; The entity segment internal token includes two parts, the first part marks that the current token belongs to the entity segment internal token; the second part marks the entity type, and the entity type includes subject and object; The other non-entity token includes a part, the first part marks that the current token belongs to the other non-entity token; The first token of the entity segment is marked as the entity segment start token, all subsequent tokens starting from the second token in the entity segment are marked as internal tokens of the entity segment, and the entity types of the corresponding tags of all tokens in the entity segment are the same.
3. The prompt-based entity-level autoregressive relationship triple extraction model training method according to claim 1 is characterized in that: The classifier includes a continuous fully connected layer and a Softmax layer; the BERT encoder is pre-trained using the text data of the target subject environment.
4. The prompt-based entity-level autoregressive relationship triple extraction model training method according to claim 3 is characterized in that: The method also includes: pruning the attention heads and layers of the BERT encoder; and introducing a local attention mechanism or a cross-segment attention mechanism to the BERT encoder.
5. The prompt-based entity-level autoregressive relationship triple extraction model training method according to claim 3 is characterized in that: The BERT encoder adopts DistilBERT / TinyBERT.
6. The prompt-based entity-level autoregressive relation triple extraction model training method according to claim 1, characterized in that: A cross entropy loss function is constructed based on the prediction result and the label after label smoothing. The expression of the loss is: Among them, y LS represents the label after smoothing, where j represents the vector y LS The index of each class in , j0 represents the index of the actual category, and α is a hyperparameter indicating the degree of smoothness.
7. The prompt-based entity-level autoregressive relation triple extraction model training method according to claim 1, characterized in that: The method also includes: dividing the training sample set into a training set, a validation set and a test set, using the training set to perform model training, using the validation set to evaluate and tune the target relationship triple extraction model, using the test set to calculate the standard accuracy, recall rate and F1 score of the trained and tuned target relationship triple extraction model, and generating a performance evaluation report.
8. A prompt-based entity-level autoregressive relation triple extraction method, characterized in that: The method comprises the following steps: Obtain sentences to be recognized related to the target subject environment; After tokenizing the sentence to be recognized, the token is input into the target relation triple extraction model in the prompt-based entity-level autoregressive relation triple extraction model training method described in any one of claims 1 to 7, and the entity recognition result is output.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Entity relationship extraction method and system for knowledge graph construction
CN113626537A
Entity relation joint extraction method and device
CN114757179A
Entity relationship extraction method and device, electronic equipment and storage medium
CN115357723A
Information extraction method and device
CN115599925A
Entity relationship extraction model training method and device and readable storage medium
CN115757811A