Entity relationship extraction method, device and equipment
By combining few-shot learning methods and entity label replacement strategies with BERT models and prototype networks, the problem of insufficient model generalization in entity relation extraction in the field of aerospace manufacturing and assembly is solved, achieving efficient entity relation extraction and structured information recognition.
Patent Information
- Application Number
- CN202511107386.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-25
AI Technical Summary
Existing deep learning-based entity relationship extraction methods are ill-suited to small sample data in the aerospace manufacturing and assembly field, resulting in insufficient model generalization ability and an inability to effectively extract complex entity relationships.
Employing a few-shot learning approach, this study uses a pre-trained BERT model and a prototype network model, combined with entity labeling and matching space replacement strategies, to label and replace text datasets in the field of aerospace manufacturing and assembly, generate input samples, and perform encoding and prototype calculations to determine entity relationship types.
It improves the model's generalization ability and the accuracy of entity relation extraction under small sample conditions, solves the problem of entity relation extraction in the field of aerospace manufacturing and assembly, and realizes efficient relation semantic recognition and structured information extraction.
Smart Images

Figure CN121009972A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of entity relationship classification technology, and in particular to an entity relationship extraction method, apparatus and equipment. Background Technology
[0002] Entity relation extraction is a key technology for effectively identifying relationships between two or more entities after named entity recognition. Relationships are the links between entities and are crucial for combining originally independent entities into a graph-based network. However, relation extraction is a rapidly developing field, and many different methods and approaches are being developed and explored, such as rule-based methods, statistical methods, traditional machine learning-based methods, and deep learning-based methods.
[0003] However, manufacturing data often contains multiple categories, with some head categories having a large number of samples while tail categories have only a very small number of samples. This situation makes some commonly used deep learning-based relation extraction methods unsuitable. For tasks with small datasets, to address the problem of insufficient long-tail relation data in relation extraction, researchers began to consider using few-shot learning to solve the long-tail relation extraction problem.
[0004] The core idea of few-shot learning is to learn a general feature representation through pre-training on a large-scale dataset, and then fine-tuning the model with a small amount of labeled data in few-shot learning to adapt it to new tasks. For example, some researchers have adopted remotely supervised few-shot relation extraction, using a global relation graph to more effectively learn new relations from different relations, employing Bayesian meta-learning to learn the posterior distribution from the prototype vectors of relations, and using graph neural networks to parameterize the initial prior distribution. Ye et al. proposed a multi-level matching aggregation grid, which encodes query samples and support set samples in an interactive manner by considering local and instance matching information, thus fully learning the associations between training samples and mining the potential information in small samples. They also learned a linear layer instead of distance for classification, which is beneficial to improving the accuracy of relation extraction. Experiments comparing the accuracy of relation extraction models using Euclidean distance and cosine distance metrics show that the metric method affects the accuracy of relation extraction, and the Euclidean distance metric model achieves higher accuracy. Soares et al. used the BERT model to encode statements containing entity relations. After testing the impact of BERT models with different input and output methods on relation extraction results, they proposed the "Matching the blanks" method to pre-train a task-agnostic relation extraction model, thus obtaining a general relation extractor that overcomes the weakness of general models in generalization.
[0005] However, entity relationship extraction in the field of aerospace manufacturing and assembly mainly focuses on the extraction of entity relationships of key features such as components (component models), parts, parameters, materials, functions, structures, and characteristics involved in web pages, documents, patents, and technical reports. Currently, there is little involvement of this both domestically and internationally. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an entity relation extraction method, apparatus and equipment to improve the model's ability to model relation semantics and generalize under small sample conditions, thereby ensuring entity relation extraction in the field of aerospace manufacturing and assembly where the entity relation structure is complex and sample acquisition is difficult.
[0007] To address the aforementioned technical problems, embodiments of the present invention provide an entity relation extraction method, comprising the following steps:
[0008] A text dataset of corpus in the field of aerospace manufacturing and assembly is obtained, and the text dataset is divided into a support set and a query set. Both the support set and the query set contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample.
[0009] The entity relationship statement samples in the support set and the query set are marked and replaced sequentially to obtain the first input sample corresponding to the support set and the second input sample corresponding to the query set; the entity replacement process is to match and replace spaces in the entities in the marked samples to construct matching and space-replaced samples.
[0010] The first input sample and the second input sample are input into the pre-trained BERT model for encoding processing to obtain the first entity relation vector in the first input sample and the second entity relation vector in the second input sample;
[0011] Based on a preset prototype network model, prototype calculation and classification processing are performed on the first entity relationship vector and the second entity relationship vector to determine the predicted target relationship type between entities in the query set.
[0012] Based on a preset loss function, the parameters of the pre-trained BERT model and the preset prototype network model are backpropagated and updated to complete end-to-end small sample entity relationship extraction training.
[0013] In one embodiment, entity relationship statement samples in the support set and the query set are sequentially marked and replaced to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set, including:
[0014] Based on preset entity markers, entity marker processing is performed on the entities in the entity relationship statements of the support set and the query set respectively to obtain the first marker sample corresponding to the support set and the second marker sample corresponding to the query set;
[0015] Based on a preset space replacement mechanism, the entities in the first marked sample and the entities in the second marked sample are replaced to obtain the first replacement sample corresponding to the first marked sample and the second replacement sample corresponding to the second marked sample.
[0016] The first input sample is determined based on the first labeled sample and the first replacement sample;
[0017] The second input sample is determined based on the second labeled sample and the second replacement sample.
[0018] In one embodiment, entity tagging processing is performed on entities in entity relationship statements of the support set and the query set based on preset entity taggers to obtain a first tag sample corresponding to the support set and a second tag sample corresponding to the query set, including:
[0019] Preset start markers and preset end markers are embedded at the beginning and end of each entity in each entity relation statement in the support set and the query set, respectively, to mark the entities.
[0020] In one embodiment, entities in the first marked sample and entities in the second marked sample are replaced based on a preset space replacement mechanism to obtain a first replacement sample corresponding to the first marked sample and a second replacement sample corresponding to the second marked sample, including:
[0021] A replacement strategy for a preset space replacement mechanism is determined, the replacement strategy including: a preset replacement probability and a preset number of entity replacements;
[0022] Based on the preset replacement probability and the preset number of entity replacements, entities in the first marked sample and entities in the second marked sample are replaced by preset space identifiers to obtain a first replacement sample corresponding to the first marked sample and a second replacement sample corresponding to the second marked sample.
[0023] In one embodiment, the first input sample and the second input sample are input into a pre-trained BERT model for encoding processing to obtain a first entity relation vector in the first input sample and a second entity relation vector in the second input sample, including:
[0024] The first input sample and the second input sample are respectively subjected to preset vector embedding processing to obtain the first input vector corresponding to each entity in the first input sample and the second input vector corresponding to each entity in the second input sample;
[0025] The first input vector and the second input vector are input into multiple Transformer encoder layers of the pre-trained BERT model for multi-level processing to obtain the first hidden vector corresponding to each entity in the first input sample and the second hidden vector corresponding to each entity in the second input sample; both the first hidden vector and the second hidden vector contain contextual and semantic information from the entity relationship statement samples.
[0026] The first hidden vectors of each entity in the first input sample are concatenated to obtain the first entity relation vector in the first input sample.
[0027] The second hidden vectors of each entity in the second input sample are concatenated to obtain the second entity relation vector in the second input sample.
[0028] In one embodiment, the first input sample and the second input sample are respectively subjected to preset vector embedding processing to obtain a first input vector corresponding to each entity in the first input sample and a second input vector corresponding to each entity in the second input sample, including:
[0029] The first input sample and the second input sample are respectively processed by word segmentation to obtain the first input code corresponding to the first input sample and the second input code corresponding to the second input sample; both the first input code and the second input code contain entity word segmentation and entity relation word segmentation.
[0030] The word segments in the first input encoding and the second input encoding are subjected to preset vector embedding processing to obtain the first input vector corresponding to the first input encoding and the second input vector corresponding to the second input encoding.
[0031] In one embodiment, prototype calculation and classification processing are performed on the first entity relationship vector and the second entity relationship vector based on a preset prototype network model to determine the predicted target relationship type between entities in the query set, including:
[0032] A prototype calculation is performed on the first entity relation vector corresponding to the entity relation statement sample in the support set using a preset prototype network model, so as to generate the prototype vector of the corresponding relation type in each entity relation statement sample in the support set.
[0033] Based on the Euclidean distance between the second entity relation vector corresponding to the entity relation statement sample in the support set and the prototype vector, the predicted target relation type between entities in the query set is determined.
[0034] Embodiments of the present invention also provide an entity relation extraction device, comprising:
[0035] The acquisition module is used to acquire a corpus text dataset in the field of aerospace manufacturing and assembly, and divide the corpus text dataset into a support set and a query set. The support set and the query set each contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample.
[0036] The processing module is used to sequentially label and replace entity relationship statement samples in the support set and the query set to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set. The entity replacement process involves matching and replacing spaces in the entities in the labeled samples to construct matching and space-replaced samples. The first input sample and the second input sample are input into a pre-trained BERT model for encoding to obtain a first entity relationship vector in the first input sample and a second entity relationship vector in the second input sample. Based on a preset prototype network model, the first entity relationship vector and the second entity relationship vector are subjected to prototype calculation and classification to determine the predicted target relationship type between entities in the query set. Based on a preset loss function, the parameters of the pre-trained BERT model and the preset prototype network model are backpropagated and updated to complete end-to-end small sample entity relationship extraction training.
[0037] Embodiments of the present invention also provide a computing device, comprising:
[0038] Memory, used to store one or more programs;
[0039] One or more processors are configured to execute the one or more programs to implement the method described above.
[0040] Embodiments of the present invention also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the method described above.
[0041] The above-described solution of the present invention has at least the following beneficial effects:
[0042] The solution provided by the above embodiments of the present invention obtains a corpus text dataset in the field of aerospace manufacturing and assembly, and divides the corpus text dataset into a support set and a query set; it then marks and replaces entity relation statement samples in the support set and the query set to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set; it inputs the first input sample and the second input sample into a pre-trained BERT model for encoding to obtain a first entity relation vector in the first input sample and a second entity relation vector in the second input sample; it performs prototype calculation and classification processing on the first entity relation vector and the second entity relation vector based on a preset prototype network model to determine the predicted target relation type between entities in the query set; and it backpropagates and updates the parameters of the pre-trained BERT model and the preset prototype network model based on a preset loss function to complete end-to-end small sample entity relation extraction training. This application leverages the unique advantages of few-shot learning and utilizes a small amount of labeled data to address the long-tail problem of Chinese entity relations. This enables the relation extraction model based on few-shot learning to exhibit more powerful and efficient learning capabilities. It solves the problems of existing relation extraction methods struggling to obtain high-quality semantic representations and declining relation classification performance when labeled samples are scarce. Consequently, it improves the accuracy and efficiency of entity relation extraction in the aerospace manufacturing and assembly field, where entity relation structures are complex and sample acquisition is difficult. Attached Figure Description
[0043] Figure 1 This is a flowchart of the entity relationship extraction method provided in the embodiments of the present invention;
[0044] Figure 2 This is a comparison chart of ablation experiment results on the AM_NER dataset provided by an optional embodiment of the present invention;
[0045] Figure 3 This is a schematic block diagram of the entity relationship extraction device provided in an embodiment of the present invention;
[0046] Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention;
[0047] Figure 5 This is a schematic block diagram of a computing device provided in an embodiment of the present invention. Detailed Implementation
[0048] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0049] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known apparatuses, structures, and techniques associated with this application may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0050] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.
[0051] like Figure 1 As shown, an embodiment of the present invention provides an entity relation extraction method, including the following steps:
[0052] Step 11: Obtain the text corpus dataset in the field of aerospace manufacturing and assembly, and divide the text corpus dataset into a support set and a query set. Both the support set and the query set contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample.
[0053] Step 12: The entity relationship statement samples in the support set and query set are marked and replaced sequentially to obtain the first input sample corresponding to the support set and the second input sample corresponding to the query set; the entity replacement process is to match and replace spaces in the entities in the marked samples to construct matching and space-replaced samples.
[0054] Step 13: Input the first input sample and the second input sample into the pre-trained BERT model for encoding processing to obtain the first entity relation vector in the first input sample and the second entity relation vector in the second input sample;
[0055] Step 14: Based on the preset prototype network model, perform prototype calculation and classification processing on the first entity relationship vector and the second entity relationship vector to determine the predicted target relationship type between entities in the query set.
[0056] Step 15: Based on the preset loss function, backpropagate and update the parameters of the pre-trained BERT model and the preset prototype network model to complete the end-to-end small sample entity relationship extraction training.
[0057] In this embodiment, the corpus text dataset contains multiple entity relation statement samples, and each entity relation statement sample contains entity pair samples (entity 1, entity 2) and their corresponding entity relation types;
[0058] After obtaining the text corpus dataset, it can be divided into a training set and a test set. Initial labeling is performed on the entity categories and relation types in the entity relation sentence samples in the training set. Taking the aerospace manufacturing and assembly domain data as an example, entity categories such as "Model", "Parameter", "Part", "Structure", "Material", "Brand", and "Principle" are labeled, along with corresponding entity relation types such as "Has_a_part", "Has_a_material", "Has_a_model", "principle_to", and "Has_a_feature". Further, a mapping from relation types to entity relations is constructed, and a support set and a query set are divided in the training set to form a support-query training strategy for subsequent model training. Simultaneously, for each type of relation in the support set and query set of the training set, a mapping table from relation type to relation instances between entity pairs is established to assist the model in quickly matching known relations in the support set during the inference phase.
[0059] Here, entity tagging and entity replacement are performed on entity relation statement samples in the support set and query set to change the input format required for training the pre-trained BERT model. This obtains first and second input samples that assist the pre-trained BERT model in semantic boundary recognition, ensuring the accuracy of subsequent model training and thus the accuracy of entity relation extraction based on the trained model. It should be noted that each entity relation statement sample corresponds to one input sample. Furthermore, the first and second input samples are input into the pre-trained BERT model for multi-level training, enabling the pre-trained BERT model to learn the mapping from relation statements to relation representations. Here, the relation representation is the entity relation vector output by the model.
[0060] Furthermore, based on the pre-set prototype network model, the second entity relation vector corresponding to the entity relation statement sample in the query set is compared with the first entity relation vector corresponding to the entity relation statement sample in the support set, and the predicted target relation type between entities in the query set is obtained based on the similarity obtained from the comparison. In the model training process of this embodiment, a pre-set loss function can be used for joint training to update the model parameters, improve the generalization performance and robustness of the pre-trained BERT model and the pre-set prototype network model in small sample scenarios, and thus effectively improve the performance of the model under a small amount of labeled data. The trained model is deployed in text understanding tasks such as the assembly of UAV power systems to realize the automatic extraction of relation information of structured entities from technical documents.
[0061] In an optional embodiment of the present invention, step 12 above may include:
[0062] Step 121: Based on the preset entity markers, perform entity marking processing on the entities in the entity relationship statements of the support set and the query set respectively, so as to obtain the first marked sample corresponding to the support set and the second marked sample corresponding to the query set.
[0063] In this embodiment, the standard input of the pre-trained BERT model is improved to construct labeled samples containing entity labels.
[0064] Specifically, step 121 above may include:
[0065] Step 1211: Embed the preset start marker and preset end marker into the beginning and end positions of each entity in each entity relation statement in the support set and query set, respectively, to mark the entities.
[0066] Here, in the original BERT model's input structure, standard paragraph start symbols [CLS] and paragraph end symbols [SEP] are first used to mark the start and segment of entity relation sentence samples, respectively, to assist the pre-trained BERT model in recognizing semantic boundaries. To facilitate the model's focus on the position of entities in the entity relation sentence input text, entity positions are explicitly labeled, and entity positions are encoded in the input text as natural numbers to enhance the model's ability to perceive entity position information. At the same time, preset start and end markers for entity pairs in the entity relation sentence samples are defined and represented as [E1start], [E1end], [E2start], and [E2end], respectively. These four markers correspond to the four types of preserved segment embeddings in the pre-trained BERT model input, used to explicitly enhance the positional information of entity boundaries.
[0067] Here, an entity relationship can be defined as follows: ,in Defined as a sequence of relations in an entity-relation statement. Defined as the beginning of a paragraph. , Defined as paragraph end symbol ;entity and entity Span identifier and It can be represented as: , Where i, j, k, and l represent the positions of entities in the statement, and are integers of type 0 < i < j-1, k < l-1, and l ≤ n. and Representing the head entity and tail entity respectively, four types of reserved phrases are further introduced: , , and Then, a tokenized sample containing the entity and the end marker can be represented as:
[0068] ;
[0069] at this time, and Simultaneously updated: , .
[0070] In an optional embodiment of the present invention, step 12 may further include:
[0071] Step 122: Based on a preset space replacement mechanism, the entities in the first marked sample and the entities in the second marked sample are replaced to obtain the first replacement sample corresponding to the first marked sample and the second replacement sample corresponding to the second marked sample.
[0072] In this embodiment, to enhance the model's ability to extract relational semantics under sparse sample conditions, a matching whitespace mechanism is introduced. This involves replacing spaces in entities in the first and second labeled samples to construct matching whitespace-replaced samples. This allows the model to learn to focus on the relational semantics themselves during training, rather than overly relying on specific entities, thereby enhancing the model's ability to learn relational semantics. It should be noted that the preset whitespace replacement mechanism performs replacement operations on entities in a subset of samples within the support set and query set, rather than replacing entities in all samples.
[0073] In an optional embodiment of the present invention, step 122 above may include:
[0074] Step 1221: Determine the replacement strategy of the preset space replacement mechanism. The replacement strategy includes: preset replacement probability and preset number of entity replacements.
[0075] Step 1222: Based on the preset replacement probability and the preset number of entity replacements, replace the entities in the first marked sample and the entities in the second marked sample with a preset space identifier to obtain the first replacement sample corresponding to the first marked sample and the second replacement sample corresponding to the second marked sample.
[0076] In this embodiment, when constructing matching space replacement samples (i.e., the first replacement sample and the second replacement sample), entities in a portion of the labeled samples (the first labeled sample and the second labeled sample) are randomly selected according to a certain probability or strategy and replaced with a preset space identifier [BLANK]. Here, a preset replacement probability p0 can be set. For each entity in each labeled sample, it is replaced with the preset space identifier [BLANK] with the preset replacement probability p0. This enriches the diversity of the training data, provides more training samples for subsequent model training, and helps the model learn a more robust relation representation.
[0077] Here, the entity in the replaced labeled sample can be the head entity, the tail entity, or both entities can be replaced simultaneously, depending on the task requirements and data characteristics. For example, for the labeled sample “[CLS] the [E1start]engine model Y-20 [E1end] has a [E2start] high-efficiency fuel injection system [E2end] [SEP]”, a preset replacement probability of 0.5 is used to determine whether to replace one or both entities; replacing the head entity results in “[CLS] the [E1start] [BLANK] [E1end] has a [E2start] high-efficiency fuel injection system [E2end] [SEP]”, replacing the tail entity results in “[CLS]the [E1start] engine model Y-20 [E1end] has a [E2start] [BLANK] [E2end][SEP]”, or replacing both entities results in “[CLS] the [E1start] [BLANK] [E1end] has a [E2start] [BLANK] [E2end] [SEP]”.
[0078] In an optional embodiment of the present invention, step 12 may further include:
[0079] Step 123: Determine the first input sample based on the first labeled sample and the first replacement sample;
[0080] Step 124: Determine the second input sample based on the second labeled sample and the second replacement sample.
[0081] In this embodiment, the constructed matching space replacement samples are integrated into the first labeled samples of the original support set and the second labeled samples of the query set to form an augmented sample dataset: the first input sample and the second input sample. These samples will be used together in the model pre-training and fine-tuning stages to train the model's understanding of relational semantics and its ability to extract entity relations.
[0082] In an optional embodiment of the present invention, step 13 above may include:
[0083] Step 131: Perform preset vector embedding processing on the first input sample and the second input sample respectively to obtain the first input vector corresponding to each entity in the first input sample and the second input vector corresponding to each entity in the second input sample.
[0084] Specifically, step 131 above may include:
[0085] Step 1311: Perform word segmentation on the first input sample and the second input sample respectively to obtain the first input code corresponding to the first input sample and the second input code corresponding to the second input sample; both the first input code and the second input code contain entity word segmentation and entity relation word segmentation;
[0086] Step 1312: Perform preset vector embedding processing on the word segments in the first input code and the second input code respectively to obtain the first input vector corresponding to the first input code and the second input vector corresponding to the second input code.
[0087] Here, the tagged sentence is segmented to meet the input requirements of the pre-trained BERT model. Preferably, the pre-trained BERT model typically uses the WordPiece segmentation algorithm to decompose the input sample into words or sub-word units. For example, the input sample "[CLS] The [E1start] drone's engine model Y-20 [E1end] has a [E2start] high-efficiency fuel injection system [E2end] [SEP]" can be represented as follows after segmentation:
[0088] "[CLS], the, [E1start], drone's, engine, model, Y-20, [E1end], has a, [E2start, high-efficiency, fuel, injection, system, [E2end], [SEP]";
[0089] Furthermore, the segmented tokens (words or sub-word units) are mapped to the vocabulary of the pre-trained BERT model, generating corresponding input identifiers (IDs). Each token has a unique input identifier (ID), and these IDs form an encoding that serves as the input encoding for the pre-trained BERT model. Further, an attention mask is created to indicate which tokens the pre-trained BERT model should focus on. For actual tokens, the mask value is 1; for padded tokens (entities replaced by the pre-defined whitespace identifier [BLANK]), the mask value is 0, so that the pre-trained BERT model can correctly ignore padded parts when processing sequences of different lengths.
[0090] Furthermore, the first and second input codes undergo pre-defined vector embedding processing to convert them into input vectors for easier training of the pre-trained BERT model. These pre-defined vector embeddings include word embedding vectors, position embedding vectors, paragraph embedding vectors, and token embedding vectors. Word embedding vectors are generated by converting the input codes into word embedding vectors. Word embeddings are distributed representations of words that capture their semantic information. The pre-trained BERT model uses a predefined word embedding matrix; input IDs are retrieved by looking up the corresponding word embedding vectors in this matrix. Position embedding vectors are added to indicate the position of each token in the sentence and capture word order information. Paragraph embedding vectors are added to distinguish different sentences. Token embedding vectors (such as [CLS], [SEP], [E1start], [E1end], [E2start], [E2end], etc.) are added to help the pre-trained BERT model understand the structure of the input samples and the location of entities.
[0091] These final embedding vectors will be used as input to the pre-trained BERT model, which will then be fed into a multi-layer Transformer encoder for further processing. Each input vector is a point in a high-dimensional space (e.g., 768-dimensional), containing the semantic information of the token and its position within the sentence. In this way, the pre-trained BERT model can understand the meaning of each token and the relationships between them, thus enabling effective entity relation extraction in subsequent steps.
[0092] In an optional embodiment of the present invention, step 13 above may further include:
[0093] Step 132: Input the first input vector and the second input vector into multiple Transformer encoder layers of the pre-trained BERT model for multi-level processing to obtain the first hidden vector corresponding to each entity in the first input sample and the second hidden vector corresponding to each entity in the second input sample; both the first hidden vector and the second hidden vector contain contextual information and semantic information from the entity relation statement samples.
[0094] In this embodiment, the first input vector and the second input vector are fed into multiple Transformer encoder layers of the pre-trained BERT model for multi-level processing. Each Transformer encoder layer consists of a multi-head self-attention mechanism and a feedforward neural network. Within each Transformer encoder layer, a multi-head self-attention mechanism is first applied, allowing the pre-trained BERT model to establish connections between tokens at different locations, capturing their contextual relationships. Further, after the multi-head self-attention mechanism, the embedding vector of each token is further processed by the feedforward neural network. The feedforward neural network performs a non-linear transformation on the embedding vector of each token, enhancing the model's expressive power.
[0095] After processing through all Transformer encoder layers, the model outputs a hidden vector for each entity in each input sample. This hidden vector contains rich contextual and semantic information, representing the role and meaning of each entity in the entity relation statement. Through these steps, the labeled samples are input into the pre-trained BERT model for encoding, obtaining a hidden vector for each token, and these hidden vectors are used to extract entity relations. This process enables the model to fully utilize contextual information and semantic features to accurately identify the relationships between entities.
[0096] In an optional embodiment of the present invention, step 13 above may further include:
[0097] Step 133: Concatenate the first hidden vectors of each entity in the first input sample to obtain the first entity relation vector in the first input sample;
[0098] Step 134: Concatenate the second hidden vectors of each entity in the second input sample to obtain the second entity relation vector in the second input sample.
[0099] In this embodiment, the hidden vectors corresponding to the entity start markers (such as [E1start], [E2start]) in the first or second input sample are extracted. These two hidden vectors are concatenated to generate a first entity relation vector or a second entity relation vector. For example, the hidden vectors corresponding to [E1start] and [E2start] are extracted and concatenated to obtain an entity relation vector for subsequent relation classification tasks; specifically, the first entity relation vector or the second entity relation vector r... h It can be represented as:
[0100] ;
[0101] in, , These are the hidden vectors representing the starting marker positions of entity 1 and entity 2, respectively.
[0102] In an optional embodiment of the present invention, step 14 above may include:
[0103] Step 141: Use a preset prototype network model to perform prototype calculation on the first entity relation vector corresponding to the entity relation statement sample in the support set, so as to generate the prototype vector of the corresponding relation type in each entity relation statement sample in the support set.
[0104] Step 142: Determine the predicted target relationship type between entities in the query set based on the Euclidean distance between the second entity relationship vector and the prototype vector corresponding to the entity relationship statement sample in the support set.
[0105] In this embodiment, the mean of the first entity relation vector corresponding to all entity sample pairs of each relation category in the support set is calculated based on a preset prototype network model to generate the prototype vector corresponding to that category.
[0106] Preferably, the category prototype of the preset prototype network model is calculated by the following formula:
[0107] ;
[0108] Where c represents the sample category; The M-dimensional prototype vector is calculated for the pre-defined prototype network model; For sample category c y Support set; This refers to the BERT hidden vector embedding function; These are learnable parameters.
[0109] Furthermore, for each entity sample pair in the query set, the distance between the second entity relation vector corresponding to that entity sample and the prototype vectors of each category is calculated. This distance can be represented by Euclidean distance, and the minimum distance criterion is used to complete the classification task; specifically, as shown below:
[0110] ;
[0111] in, Let d be the query sample. With the prototype vector of category c Distance between Here, the model categorizes the query sample into the prototype class with the smallest distance, and uses standard cross-entropy loss to train the model, with the goal of minimizing the distance difference between the query sample and the prototype of its class.
[0112] Preferably, to determine whether two relation entities belong to the same relation type, a binary classifier based on sentence vector similarity is designed, as shown in the following formula:
[0113] ;
[0114] Where p is the similarity probability. For relational statement mapping functions, this formula represents assigning probabilities to... and Coding similarity Or not encoded The situation.
[0115] In an optional embodiment of the present invention, step 15 may include:
[0116] Step 151: Construct a preset loss function to minimize the prediction error. The preset loss function is shown in the following equation:
[0117] ;
[0118] in, Indicates using two entities and A corpus of tagged relational statements, and Each item in the relation statement is generated by... and respectively correspond to span and Two entities and Created by pairing It is a tensor product function. The value can be expressed as follows:
[0119] .
[0120] The embodiments of the present invention effectively improve the accuracy and generalization ability of relation modeling by introducing a matching space mechanism and entity annotation position enhancement model training. In the relation extraction task, a Prototypical Networks model is introduced to complete few-shot classification. This joint strategy not only fully leverages the advantages of unsupervised pre-training but also significantly reduces the dependence on large amounts of labeled data. The method provided by the above embodiments is widely applicable to highly specialized scenarios with low annotation resources, such as industry and aerospace. It exhibits excellent few-shot generalization ability, especially in specialized fields such as UAV power system assembly, providing strong support for the efficient and automated identification of related entity relationships. Simultaneously, the model structure is simple, the training process is efficient, it is compatible with the existing BERT architecture, and it is easy to deploy and implement on existing platforms.
[0121] To verify the effectiveness of the entity relation extraction method proposed in the above embodiments of the present invention, the solution provided in the above embodiments will be described below in conjunction with two specific experiments.
[0122] Example 1: Ablation Experiment
[0123] This embodiment primarily analyzes the sensitivity and robustness of the proposed method during the training and optimization phase by comparing model performance under different hyperparameter settings, thereby determining the optimal parameter configuration. The corpus used in the experiment is the self-built dataset AM_NER. The experiment includes the following steps:
[0124] Variable settings: Select key hyperparameters that affect model performance, mainly including three dimensions: optimization algorithm (such as Adam, SGD, AdamW, etc.), learning rate, and batch size.
[0125] Parameter combination construction: For each hyperparameter, multiple candidate values are set to form multiple experimental groups. For example, the learning rate is set to {1e-1, 1e-2, 1e-3, 1e-3}, and the batch size is set to {2, 4, 6}, etc.
[0126] Model training and evaluation: The proposed model is trained under each set of parameter configurations and its performance is evaluated on the validation set using the standard metric F1-score.
[0127] Statistical analysis of experimental results: The experimental results are as follows Figure 2 As shown in Table 1.
[0128] Table 1 shows the time consumption statistics for each model on the AM_NER dataset.
[0129]
[0130] Experimental results show that the optimal model is obtained, with SGD as the optimization function, learning rate lr=1e-1, and batch size m=4. The results verify the influence of each parameter on the model performance and provide theoretical support for the model's generalization.
[0131] Example 2: Comparative Experiment
[0132] After completing the ablation experiments and determining the optimal model parameter configuration, comparative experiments were conducted to further verify the superiority of the method of this invention over existing technologies. These experiments were performed on multiple publicly available datasets, including AM_NER, SemEvalTask8, and FewRel.
[0133] The specific experimental steps are as follows:
[0134] Baseline model selection: Four relation extraction models were selected as comparison baselines, including: Siamese, Prototypical networks, Snail, and FGN.
[0135] Table 2 shows the comparative experimental results of our proposed method and the baseline model on the AM_NER and SemEval task8 datasets.
[0136]
[0137] Unified configuration and training: Train each model separately under the same training set partition and experimental environment, and evaluate the results on the test set.
[0138] Performance metrics comparison: Calculate and compare the standard evaluation metric F1-score for each model on the three datasets mentioned above.
[0139] Table 2 shows the test results of the baseline model and the proposed method on the AM_NER and SemEval task8 datasets. It can be seen that, compared with the baseline model, the proposed method performs excellently in various test scenarios on the AM_NER dataset. In particular, the proposed method achieves the best F1 scores in the 5-Way-3-Shot and 5-Way-5-Shot tests, at 72.28% and 76.19%, respectively. On the SemEval task8 dataset, the proposed method achieves good F1 scores in the 5-Way-1-Shot, 5-Way-3-Shot, and 5-Way-5-Shot tests, at 82.82%, 88.68%, and 91.90%, respectively. These results demonstrate that the proposed model effectively extracts key semantic features in the relation extraction task of UAV power system assembly data.
[0140] To further verify the effectiveness of the proposed method, it was applied to the finely labeled relation extraction dataset Fewrel. This dataset is applicable not only to classic supervised / distant supervised relation extraction tasks but also shows great promise for few-shot learning tasks. The test results of the proposed method on the FewRel_procsess dataset are shown in Table 3.
[0141] Table 3 shows the results of the method presented in this paper in FewRel_procsess.
[0142]
[0143] Comparative experiments verified the advanced nature and versatility of the proposed method. In particular, in complex relationship extraction scenarios with limited labeled data and high semantic understanding requirements, the proposed method has stronger representation capabilities and better performance.
[0144] like Figure 3 As shown, embodiments of the present invention also provide an entity relationship extraction device 30, comprising:
[0145] The acquisition module 31 is used to acquire a corpus text dataset in the field of aerospace manufacturing and assembly, and divide the corpus text dataset into a support set and a query set. The support set and the query set each contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample.
[0146] Processing module 32 is used to sequentially mark and replace entity relationship statement samples in the support set and the query set to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set; the entity replacement process is to match and replace spaces in the entities in the marked samples to construct matching and space-replaced samples; the first input sample and the second input sample are input into a pre-trained BERT model for encoding to obtain a first entity relationship vector in the first input sample and a second entity relationship vector in the second input sample; the first entity relationship vector and the second entity relationship vector are prototype calculated and classified based on a preset prototype network model to determine the predicted target relationship type between entities in the query set; the parameters of the pre-trained BERT model and the preset prototype network model are backpropagated and updated based on a preset loss function to complete end-to-end small sample entity relationship extraction training.
[0147] It should be noted that this device is a device corresponding to the above-described entity relationship extraction method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0148] like Figure 4 As shown, embodiments of the present invention also provide an electronic device 50, including: a memory 51 for storing one or more computer programs; and one or more processors 52 for executing the one or more computer programs. When the computer programs are run by the processors, they perform the entity relation extraction method as described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects. The electronic device 50 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown in this invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0149] like Figure 5 As shown, electronic device 50 is a computing device or computer system, which may include CPU 501 (computing unit), which can perform various appropriate actions and processes according to a computer program stored in ROM 502 (read-only memory) or a computer program loaded from storage unit 508 into random access RAM 503 (memory). RAM 503 may also store various programs and data required for the operation of device 500. CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O interface 505 (input / output interface) is also connected to bus 504.
[0150] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0151] CPU 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of CPU 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. CPU 501 performs the various methods and processes described above. For example, in some embodiments, the entity relation extraction method can be implemented as a computer software program tangibly contained in a computer-readable storage medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by CPU 501, one or more steps of the entity relation extraction method described above can be performed. Alternatively, in other embodiments, CPU 501 can be configured to perform the entity relation extraction method by any other suitable means (e.g., by means of firmware).
[0152] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the entity relation extraction method as described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0153] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0154] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0155] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0156] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0158] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0159] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of the present invention.
[0160] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code for implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps for performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.
[0161] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for extracting entity relations, characterized in that, Includes the following steps: A text dataset of corpus in the field of aerospace manufacturing and assembly is obtained, and the text dataset is divided into a support set and a query set. Both the support set and the query set contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample. The entity relationship statement samples in the support set and the query set are marked and replaced sequentially to obtain the first input sample corresponding to the support set and the second input sample corresponding to the query set; the replacement process is to match and replace spaces in the entities in the marked samples to construct matching and space-replaced samples. The first input sample and the second input sample are input into the pre-trained BERT model for encoding processing to obtain the first entity relation vector in the first input sample and the second entity relation vector in the second input sample; Based on a preset prototype network model, prototype calculation and classification processing are performed on the first entity relationship vector and the second entity relationship vector to determine the predicted target relationship type between entities in the query set. Based on a preset loss function, the parameters of the pre-trained BERT model and the preset prototype network model are backpropagated and updated to complete end-to-end small sample entity relationship extraction training.
2. The entity relation extraction method according to claim 1, characterized in that, The entity relationship statement samples in the support set and the query set are sequentially marked and replaced to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set, including: Based on preset entity markers, entity marker processing is performed on the entities in the entity relationship statements of the support set and the query set respectively to obtain the first marker sample corresponding to the support set and the second marker sample corresponding to the query set; Based on a preset space replacement mechanism, the entities in the first marked sample and the entities in the second marked sample are replaced to obtain the first replacement sample corresponding to the first marked sample and the second replacement sample corresponding to the second marked sample. The first input sample is determined based on the first labeled sample and the first replacement sample; The second input sample is determined based on the second labeled sample and the second replacement sample.
3. The entity relation extraction method according to claim 2, characterized in that, Entity tagging processing is performed on entities in the entity relationship statements of the support set and the query set based on preset entity taggers to obtain a first tag sample corresponding to the support set and a second tag sample corresponding to the query set, including: Preset start markers and preset end markers are embedded at the beginning and end of each entity in each entity relation statement in the support set and the query set, respectively, to mark the entities.
4. The entity relation extraction method according to claim 2, characterized in that, Based on a preset space replacement mechanism, entities in the first marked sample and entities in the second marked sample are replaced to obtain a first replacement sample corresponding to the first marked sample and a second replacement sample corresponding to the second marked sample, including: A replacement strategy for a preset space replacement mechanism is determined, the replacement strategy including: a preset replacement probability and a preset number of entity replacements; Based on the preset replacement probability and the preset number of entity replacements, entities in the first marked sample and entities in the second marked sample are replaced by preset space identifiers to obtain a first replacement sample corresponding to the first marked sample and a second replacement sample corresponding to the second marked sample.
5. The entity relation extraction method according to claim 3, characterized in that, The first input sample and the second input sample are input into a pre-trained BERT model for encoding processing to obtain a first entity relation vector from the first input sample and a second entity relation vector from the second input sample, including: The first input sample and the second input sample are respectively subjected to preset vector embedding processing to obtain the first input vector corresponding to each entity in the first input sample and the second input vector corresponding to each entity in the second input sample; The first input vector and the second input vector are input into multiple Transformer encoder layers of the pre-trained BERT model for multi-level processing to obtain the first hidden vector corresponding to each entity in the first input sample and the second hidden vector corresponding to each entity in the second input sample; both the first hidden vector and the second hidden vector contain contextual and semantic information from the entity relationship statement samples. The first hidden vectors of each entity in the first input sample are concatenated to obtain the first entity relation vector in the first input sample. The second hidden vectors of each entity in the second input sample are concatenated to obtain the second entity relation vector in the second input sample.
6. The entity relation extraction method according to claim 5, characterized in that, The first input sample and the second input sample are respectively subjected to preset vector embedding processing to obtain the first input vector corresponding to each entity in the first input sample and the second input vector corresponding to each entity in the second input sample, including: The first input sample and the second input sample are respectively processed by word segmentation to obtain the first input code corresponding to the first input sample and the second input code corresponding to the second input sample; both the first input code and the second input code contain entity word segmentation and entity relation word segmentation. The word segments in the first input encoding and the second input encoding are respectively subjected to preset vector embedding processing to obtain the first input vector corresponding to the first input encoding and the second input vector corresponding to the second input encoding.
7. The entity relation extraction method according to claim 1, characterized in that, Based on a preset prototype network model, prototype calculation and classification processing are performed on the first entity relationship vector and the second entity relationship vector to determine the predicted target relationship type between entities in the query set, including: A prototype calculation is performed on the first entity relation vector corresponding to the entity relation statement sample in the support set using a preset prototype network model, so as to generate the prototype vector of the corresponding relation type in each entity relation statement sample in the support set. Based on the Euclidean distance between the second entity relation vector corresponding to the entity relation statement sample in the support set and the prototype vector, the predicted target relation type between entities in the query set is determined.
8. An entity relation extraction device, characterized in that, include: The acquisition module is used to acquire a corpus text dataset in the field of aerospace manufacturing and assembly, and divide the corpus text dataset into a support set and a query set. The support set and the query set each contain multiple entity relation statement samples and multiple entity relation types. Each type contains at least one entity pair sample or entity sample. The processing module is used to sequentially mark and replace entity relationship statement samples in the support set and the query set to obtain a first input sample corresponding to the support set and a second input sample corresponding to the query set. The entity replacement process involves matching and replacing spaces in the entities in the labeled samples to construct matching and space-replaced samples; the first input sample and the second input sample are input into the pre-trained BERT model for encoding to obtain the first entity relation vector in the first input sample and the second entity relation vector in the second input sample. Based on a preset prototype network model, prototype calculation and classification processing are performed on the first entity relationship vector and the second entity relationship vector to determine the predicted target relationship type between entities in the query set. Based on a preset loss function, the parameters of the pre-trained BERT model and the preset prototype network model are backpropagated and updated to complete end-to-end small sample entity relationship extraction training.
9. A computing device, characterized in that, include: Memory, used to store one or more programs; One or more processors are configured to execute the one or more programs to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Text statement processing method and device, computer equipment and storage medium
CN111950269A
Relationship extraction method based on domain adaptation and small sample learning
CN114564960A
Traffic accident named entity recognition method based on pre-training model
CN116432645A