Unsupervised Open Relation Extraction Text Processing Method Based on Size Model Interaction

By introducing an interactive mechanism between large language model and small language model in open relationship extraction, combined with an entropy-driven feedback mechanism, the problem of low accuracy of relationship extraction in the existing technology is solved, and higher relationship recognition accuracy and model stability are achieved, and it is suitable for complex knowledge graphs and question-and-answer system scenarios.

CN119599015BActive Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411639760.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-06-13
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

The existing open relationship extraction method faces data sparsity, diversity, noise and uncertainty, resulting in low accuracy of relationship extraction and lack of unified standards and specifications, which increases the difficulty of extraction.

Method used

An unsupervised open relationship extraction method based on interaction of size and size models is adopted to generate preliminary pseudo-relationship labels through large language models, cluster and refine small language models, and introduce an entropy-driven feedback mechanism to optimize label quality through a guiding demonstration selection mechanism.

Benefits of technology

It improves the accuracy of identifying entity relationships in text data, enhances the scalability and stability of the model, is suitable for large-scale and diversified data scenarios, and supports applications such as knowledge graph construction and question-and-answer systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119599015B_ABST
    Figure CN119599015B_ABST
Patent Text Reader

Abstract

An unsupervised open relation extraction text processing method based on large and small model interaction. First, the large language model generates preliminary pseudo-relation labels from the input statements, providing important clues for subsequent feature relation clustering. Second, these pseudo-labels are used for the supervised fine-tuning of the small language model. Through the feature learning of relation instances and the relation-guided feedback algorithm, the probability distribution of each category is generated and the pseudo-labels are optimized. Finally, samples with high-confidence probabilities are selected as demonstration examples to provide positive feedback to the large language model to reprocess those instances with uncertain relations and improve relation extraction. The InstructORE model outperforms existing unsupervised baseline models and has comparable relation extraction performance to weakly supervised methods in most cases. The present invention provides strong support for applications such as constructing knowledge graphs and knowledge answering, and has broad application prospects and promotion value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and natural language processing, and specifically to an unsupervised open relation extraction text processing method based on the interaction between large and small models. Background Art

[0002] Relation extraction aims to identify semantic relations between entities from massive unstructured texts or multi-modal data, so as to provide knowledge triples for knowledge graph construction. Traditional relation extraction is mainly limited to closed and specific domains. Although excellent performance has been achieved on fixed corpora, the variability of relations and the uncertainty of domains in reality make it difficult for existing relation extraction methods to handle real-world application scenarios. In addition, in the face of the emergence of large-scale text data, it is necessary to automatically extract unknown relations from massive texts. Therefore, the research on relation extraction technology for the open domain is particularly important.

[0003] Relation extraction for the open domain can achieve the extraction of entity relations or knowledge triples without a predefined relation set, which is also a common relation extraction scenario in reality. Currently, there are mainly two types of relation extraction methods for the open domain: one is to mine new relation triples <entity 1, relation, entity 2> through text and entities, that is, open relation extraction; the other is open triple extraction. The present invention will conduct research on open relation extraction.

[0004] Early open relation extraction methods were mainly based on pattern recognition and syntactic analysis, and these methods relied on predefined patterns or rules to identify relation triples. With the development of machine learning technology, researchers began to explore machine learning-based open relation extraction methods, which use labeled training data for supervised learning to learn how to identify and extract relations from texts. Recently, the application of transfer learning and pre-trained models in open relation extraction has also increased. These methods use large-scale text data for pre-training and then perform fine-tuning or transfer learning on specific tasks to improve the performance of relation extraction. As a key link in information extraction, open relation extraction not only has important research value but also has broad application prospects, and it is applied in aspects such as knowledge graph construction, information retrieval, question answering systems, and natural language understanding.

[0005] Although research on open relation extraction has achieved certain development at home and abroad, the existing open relation extraction methods mainly rely on unsupervised self-regulation methods of models or weakly supervised learning that depend on additional training data. Therefore, many challenges still remain. The main problems faced by open relation extraction are the sparsity and diversity of data. In the text data that needs to be processed in open relation extraction, the occurrence frequency of specific relations may be very low, making the model training data sparse and difficult to capture sufficient features to accurately identify relations, resulting in lower relation extraction accuracy. In practical applications, text data often contains a large amount of noise and uncertainty, which will reduce the accuracy of the model. At the same time, the forms of relation expression are also diverse, making it more difficult to extract new relations. Secondly, due to the lack of unified standards and specifications, the relation expressions between different data sources may vary, which increases the extraction difficulty. In addition, relation representation also needs to consider the complex interactions and dependencies between entities, which makes relation representation more complex. Some previous studies have achieved good results in new relation discovery using transfer learning methods, but a large amount of training data sets are often scarce and costly in actual real-world scenarios, and high-quality training data requires a reliable knowledge base or manual annotation.

[0006] The differences compared with the prior art are as follows:

[0007] Technical comparison with the patent CN113011189A "Method, device, equipment and storage medium for extracting open entity relations"

[0008] The structure of the patent CN113011189A is mainly based on a single pre-trained generation model, such as UniLM or GPT, etc.; the core innovation of this patent lies in introducing a collaborative mechanism of large language models (such as GPT-4, etc.) and small language models (such as BERT, etc.), and achieving unsupervised deep optimization through preliminary label generation, adaptive clustering and guiding feedback in stages. There are essential differences in the overall model structure design between the two.

[0009] The unsupervised generation model of the patent CN113011189A does not have subsequent feedback optimization after completing the preliminary relation extraction; the innovation of this patent lies in introducing an entropy-driven feedback mechanism to screen labels by confidence, incorporating samples with high confidence into the "demonstration sample pool", and using these high-confidence samples to give positive feedback to the large language model to correct the generation errors of the model for low-confidence labels. There are essential differences in the model feedback mechanism between the two.

[0010] Patent CN113011189A can cope with the needs of extracting open relations within a certain range, but because the model does not have multi-layer optimization capabilities, its scalability and stability may be limited in large-scale data scenarios; this patent provides multi-stage optimization and adaptive feedback mechanisms to enable it to cope with large-scale and diversified data, and is suitable for building complex knowledge graphs, question-answering systems, natural language understanding and other open domain scenarios. The two are essentially different in applicable scenarios. Summary of the invention

[0011] To solve the above technical problems, the present invention proposes an unsupervised open relation extraction text processing method based on large and small model interaction, which aims to accurately identify the semantic associations between entities from large-scale unstructured text data to support applications such as knowledge graph construction and knowledge question answering. First, in the preliminary relationship generation stage, a large language model (such as GPT-4, etc.) is used to receive input sentences and entity pairs, form instructions through the designed prompt template, and then use these instructions to generate preliminary pseudo-relationship labels and entity pair labels. Next, in the relationship recognition optimization stage, the entity types and relationship labels generated in the first stage are clustered and refined, and a small language model (such as BERT, etc.) is trained in a self-supervised manner to obtain the fine-grained probability distribution of the output label of each input sentence on all current clustering categories. Finally, in the instructive relationship demonstration feedback stage, the present invention introduces a relatively complex instructive demonstration selection mechanism, which explicitly places samples with clearer probabilities (i.e., low entropy) in the sample demonstration pool, which provides positive feedback to the large language model, forcing the large language model to reprocess those unclear samples to provide more accurate pseudo-relationship labels. This paper cleverly combines the advantages of large pre-trained language models (such as GPT-4, etc.) and small language models (such as BERT, etc.). Through three-stage iterative optimization, it realizes the relationship extraction task in unsupervised scenarios and improves the accuracy of identifying entity relationships in text data.

[0012] To achieve the above object, the technical solution adopted by the present invention is:

[0013] The unsupervised open relation extraction text processing method based on large and small model interaction is characterized by comprising the following steps:

[0014] 1) Preliminary relation generation: Input the sentences and corresponding entities into the large language model to generate preliminary pseudo-relation labels;

[0015] Using a pre-trained large language model to receive input sentences and entity pairs (h i ,t i ), form instructions c through the designed prompt template i ; These instructions are then used to prompt a large language model to generate preliminary pseudo relation labels and entity pair labels

[0016] 2) Relationship recognition optimization: Cluster and refine the entity types and relationship labels generated in the first stage, and train a small language model in a self-supervised manner to obtain the fine-grained probability distribution of the output labels of each input sentence over all current clustering categories;

[0017] Use a pre-trained small language model to encode the text descriptions of entities to obtain preliminary pseudo-labels In the original sentence s i The context feature f i ; Cluster the obtained preliminary feature set through an adaptive clustering strategy, and use this clustering category as the training set to fine-tune the small language model to obtain refined pseudo-labels l i ; Use these optimized pseudo-labels as self-supervised signals to further fine-tune the small language model;

[0018] 3) Guided relationship demonstration feedback: Select samples with high confidence probabilities as demonstration examples to provide positive feedback to the large language model for reprocessing instances with uncertain relationships to improve the accuracy and quality of the labels;

[0019] Use the fine-tuned small language model to generate a probability distribution p i , h i , t i ) over all the pseudo-labels l i of each sample (s i , and then based on an entropy-driven demonstration example selection mechanism, use entropy to measure the uncertainty level of each instance, and select samples with a high confidence probability distribution, that is, samples with an entropy value H(p i ) below the threshold as demonstration examples. These samples are placed in the sample demonstration pool to guide the large language model to reprocess those instances with uncertain relationships to provide more accurate pseudo-relationship labels.

[0020] As a further improvement of the present invention, large language models including GPT-4, PaLM 2, and LLaMA 2 are used, and small language models including BERT, MiniLM, and ALBERT are used for interaction.

[0021] As a further improvement of the present invention, step 1) preliminary relationship generation is as follows:

[0022] Definition 1 Closed relationship extraction: For a given set of sentences Predefined relationship set Extract the specific relationship r of the head entity h i and the tail entity t i in each sentence s i in the relationship set Rk , where N s represents the number of statements in the text set, and N r represents the number of relationships in the relationship set, where 1 ≤ k ≤ N r ;

[0023] Definition 2: Open relation extraction: Open relation extraction means, without giving a predefined relation set R, finding new relations between the head entity h and the tail entity t i from the given text corpus. This task is formalized as a clustering task that groups all the relationship sets i corresponding to the current sentence set S into specific categories. The output of open relation extraction is defined as where K is the number of categories of the defined relation clustering, and the relations of each text sentence instance will be placed in one of the clustering categories,

[0024] where k ∈ [1, K]; k ∈ [1, K];

[0025] For an input sentence s i and the entity pair h i , t i in the corresponding sentence, design a prompt template and fill the position slots of the template with s i , h i and t i to form an instruction c i , and then use these instructions for a large language model to generate preliminary pseudo-relation labels and entity pair labels as shown in formula (1):

[0026]

[0027] and generate and entity labels of

[0028] As a further improvement of the present invention, step 2) Relationship recognition optimization is as follows:

[0029] After inputting the prompt instruction text c i into the large language model and obtaining the preliminary generation result , use a small language model to obtain the context features of the preliminary pseudo-labels in the original sentence. Subsequently, identify refined soft pseudo-labels through a specific clustering method and further use them to fine-tune the small language model to generate the probability of each label;

[0030] By merging the special tokens [E start and [Eend to update the original input sentence s i , representing the corresponding head entity and the tail entity 's boundary positions. In addition, [R start and [R end are included in the relation label . The purpose of introducing these special tokens is to obtain richer features of all available information such as text and labels in the small language model, the original sentence s i combined with the generated relation label is updated. Formally, after the sentence is optimized it is represented as shown in formula (2):

[0031]

[0032] For the optimized 's relation features, syntactic dependency parsing is performed on the original input sentence to extract the dependency path between the head entity and the tail entity. Further, an average pooling algorithm is applied to all the tokens on the obtained dependency path to generate the dependency feature f dep ∈ R d , as shown in formula (3):

[0033]

[0034] where represents the hidden features between the special tokens [E start , [E end , [R start , [R end . In addition, MP(.) and AP(.) represent the max pooling function and the average pooling function, represents the dependency path of length d, and in (s i , h i , t i ) the clustering feature f i ∈ R d is formed by concatenating the features of the entity, relation, and dependency, which enables f i to consider both semantic and syntactic information simultaneously, as shown in formula (4):

[0035]

[0036] where represents the concatenation symbol between the respective hidden features;

[0037] Use the small language model to obtain the preliminary feature set After that, the purpose of the pseudo-relationship label clustering is to group F into K clustering categories, where K is a fixed value or manually defined, and an adaptive clustering strategy is adopted.

[0038] As a further improvement of the present invention, the adaptive clustering strategy includes two steps:

[0039] (1) Use a feature transformation method that transforms from a high-dimensional feature representation f i ∈R d to a low-dimensional feature representation. Here, a non-linear mapping network g φ is used, where is the transformed dimension;

[0040] (2) Learn a soft assignment to assign all n instances to K centroids to obtain a better clustering effect. This strategy promotes high-confidence assignments and is insensitive to the predefined number of clusters;

[0041] First, apply k-means to obtain K initial centroids Subsequently, using the t-student distribution, calculate the feature similarity between each instance feature and the centroid , and then calculate the soft assignment label q i of each sample s ik , as shown in formula (5):

[0042]

[0043] where α is a hyperparameter used to adjust the clustering algorithm, and the frequency statistics of the entire feature set are used to be able to make an alternative estimate of the sample distribution p ik , as shown in formula (6):

[0044]

[0045] where is the normalization factor. To adaptively improve the spatial proximity of related entity pairs, use the KL divergence between the soft assignment and the auxiliary distribution to optimize the non-linear network g φ , and generate a finely optimized pseudo-label l i for each sample s i . Formally, the loss of soft clustering and the pseudo-label l i are as shown in formula (7):

[0046]

[0047] Use these optimized pseudo-labels As a self-supervised signal, further adjust the small language model to obtain a better SLM θ , with a relational classification loss As shown in formula (8):

[0048]

[0049] where is the cross-entropy loss, onehot(.) returns a one-hot vector for the pseudo-label, and the sentence, head entity, and tail entity are concatenated as the input x i , to optimize the SLM θ .

[0050] As a further improvement of the present invention, step 3) guided relational demonstration feedback is as follows:

[0051] At this stage, use the adjusted SLM θ , and generate a probability distribution p over all the pseudo-labels of each sample (s i , h i , t i ). The higher the probability value in a certain relational category, the more definite the small model's relational selection for that instance. Then use the entropy H(p i ) to measure the uncertainty level p of each instance i , as shown in formula (9): i

[0052]

[0053] where a small entropy indicates high certainty, and conversely, a large entropy indicates high ambiguity.

[0054] Compared with the prior art, the beneficial effects of the present invention are:

[0055] ​The present invention proposes an unsupervised open relation extraction text processing method based on the interaction between large and small models, aiming to accurately identify the semantic associations between entities from large-scale unstructured text data. The present invention believes that in open relation extraction, some human annotation and selection operations can be replaced by large language models. The rich prior knowledge of large language models (such as GPT-4, etc.) helps to generate robust and reliable pseudo-labels, which can be used to fine-tune small language models. Conversely, small language models (such as BERT, etc.) can enhance the capabilities of large language models by selecting more guiding demonstration examples. This continuous interaction between large language models and small language models realizes the extraction of effective data while enhancing their respective capabilities. Secondly, the model gradually improves the accuracy and reliability of relation extraction through three stages of iterative optimization. This iterative mechanism enables the model to continuously learn and adapt to new data, improving the extraction performance. Finally, the present invention introduces a complex guiding demonstration selection mechanism. By selecting samples with high confidence probabilities as demonstration examples, positive feedback is provided for the large language model to reprocess instances with uncertain relationships, improving the accuracy and quality of the labels. The InstructORE model outperforms existing unsupervised baseline models and has comparable relation extraction performance to weakly supervised methods in most cases. The present invention provides strong support for applications such as constructing knowledge graphs and knowledge answering, and has broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the logical flowchart of the present invention;

[0057] Figure 2 is the model process sub- Figure 1 ;

[0058] Figure 3 is the model process sub- Figure 2 ;

[0059] Figure 4 is the model process sub- Figure 3 . DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:

[0061] The present invention proposes an open relation extraction model InstructORE based on the interaction between large and small models, aiming to accurately identify the semantic associations between entities from large-scale unstructured text data to support applications such as knowledge graph construction and knowledge Q&A. First, in the preliminary relation generation stage, a large language model (such as GPT-4, etc.) receives the input sentence and entity pair, forms instructions through the designed prompt template, and then uses these instructions to generate preliminary pseudo-relation labels and entity pair labels. Next, in the relation recognition optimization stage, the entity types and relation labels generated in the first stage are clustered and refined, and a small language model (such as BERT, etc.) is trained in a self-supervised manner to obtain the fine-grained probability distribution of the output labels of each input sentence over all current clustering categories. Finally, in the guided relation demonstration feedback stage, the present invention introduces a relatively complex guided demonstration selection mechanism, places the samples with more explicit probabilities (i.e., low entropy) into the sample demonstration pool, which provides positive feedback to the large language model, forcing the large language model to reprocess those ambiguous samples to provide more accurate pseudo-relation labels. The present invention cleverly combines the advantages of large pre-trained language models (such as GPT-4, etc.) and small language models (such as BERT, etc.), and through three-stage iterative optimization, realizes the relation extraction task in an unsupervised scenario, improving the accuracy of identifying entity relations in text data.

[0062] As a specific embodiment of the present invention, the present invention provides a logic flow chart as Figure 1 shown, and a model flow chart as Figure 2 Figure 3 Figure 4 An open relation extraction model InstructORE based on the interaction between large and small models as shown, including the steps:

[0063] 1) Preliminary relation generation:

[0064] Definition 1 (Closed relation extraction): For a given set of sentences Pre-defined relation set Extract for each sentence s i the head entity h i and the tail entity t i in the specific relation r in the relation set R k , where N s represents the number of sentences in the text set, N r represents the number of relations in the relation set, 1 ≤ k ≤ N r .

[0065] Definition 2 (Open relation extraction): Open relation extraction is to discover the head entity h from the given text corpus without giving the pre-defined relation set R i and the tail entity ti new relationships. This task is formalized as a clustering task that groups all the relationship sets corresponding to the current sentence set S into specific categories. The output of open relation extraction is defined as where K is the number of categories of the defined relation clusters. The relationship of each text sentence instance will be placed in one of the clustering categories, for example k ∈ [1, K].

[0066] Some previous weakly supervised methods combined pre-training with pseudo-labels and manual intervention, thus achieving the best new relation extraction results for weakly supervised open relation extraction models. The present invention assumes that the rich real-world knowledge embedded in large language models can also help generate initial pseudo-labels. Such as Figure 2 , in this embodiment, GPT-4 is used as the large language model in the first stage. GPT-4 (Generative Pre-trained Transformer 4) is the latest GPT series model released by OpenAI and is currently the most powerful text generation model. It is a large-scale multi-modal model that uses a very wide variety of real-world corpora and complex semantic tasks on the Internet during the pre-training stage, can accept image and text inputs, and produce text outputs. The text output task is still an autoregressive word prediction task. The present invention uses the powerful generation ability of GPT-4 in the first stage to provide high-value pseudo-labels for the open relation extraction task.

[0067] For an input sentence s i and the entity pair h in the corresponding sentence i , t i , the present invention designs a prompt template and fills the position slots of the template with s i , h i and t i to form an instruction c i . Then these instructions are used to prompt GPT-4 to generate initial pseudo-relation labels and entity pair labels as shown in formula (1):

[0068]

[0069] For example, in a text "Two years ago, Stephen Stone organized a grand wedding for his beloved Linda Skyler in the romantic city of Paris to express deep love and commitment", Stephen Stone and Linda Skyler are the head entity and the tail entity of this sentence. The prompt texts constructed based on the text and the entities are in a continuous writing format. For example: "There is a sentence 'Two years ago... in Paris', the entity types of Stephen Stone and Linda Skyler are XXX and XXX. The relationship between Stephen Stone and Linda Skyler is XXX". To ensure that the output of the large language model is more focused on a certain relationship semantics, the present invention also makes a restriction on word counting, adding the prompt "Please answer no more than three words" in the whole prompt text string and then inputting it into the large language model GPT-4. In a similar situation, GPT-4 can also generate and entity labels. In the case of the input constructed by the prompt in the above example, the preliminary relationship generated from the output of GPT-4 is "spouse", and the entity labels and are both "person".

[0070] In addition, during the first iteration between the large language model and the small language model, the present invention adopts a zero-shot method to input the prompt set into GPT-4 to directly generate preliminary pseudo-labels. As the subsequent iteration process unfolds, the third-stage guided relationship demonstration feedback can help the overall model generate more robust pseudo-labels. Then, a method of adding a small number of demonstration samples is used to construct the prompt text string.

[0071] 2) Relationship Recognition Optimization

[0072] In this embodiment, BERT is used as the small language model in the second stage. This stage adopts the Encoder part of Transformer and is an auto-encoding language model. Generally, pre-trained language models are pre-trained from a large amount of unlabeled data, mainly to obtain more general semantic representations, thereby improving the performance of downstream tasks; secondly, it can provide better initial parameters for the model, so as to have better generalization performance and accelerate convergence. At present, BERT is mainly applied to downstream tasks in two ways, one is the Feature-extract method and the other is the Fine-tune method. The present invention adopts the fine-tuning method, which can utilize the weights of the BERT model to initialize the network and continuously adjust all the weight parameters of the model during the training process, and finally train a model that can be better applied to specific tasks. The present invention uses BERT to encode and represent the text description of entities, and performs semantic fusion by combining the context encoding representations of different Transformer blocks. The entity description encoder used in the present invention mainly includes three parts: an input layer, an encoding layer, and a mapping layer. The number of parameters of BERT is of a very different order of magnitude compared to GPT-4, and it can achieve excellent selection performance through iteration in a certain task by fine-tuning, so it is selected as the small language model for optimizing relationships in the open relation extraction model InstructORE proposed by the present invention.

[0073] As Figure 3 , after inputting the prompt instruction text c i into GPT-4 and obtaining the initial generation result , use the BERT model to obtain the context features of the initial pseudo-label in the original sentence. Subsequently, refined soft pseudo-labels are identified through a specific clustering method and further used to fine-tune BERT to generate the probability of each label. The process and algorithm details of this stage will be elaborated below.

[0074] By merging the special tokens [E start and [E end to update the original input sentence s i , which represent the boundary positions of the corresponding head entity and tail entity . In addition, [R start and [R end are included in the relation label . The purpose of introducing these special tokens is to obtain richer features of all available information such as text and labels in BERT. For example, in Figure 2 , the original sentence s i is combined with the generated relation label After being updated to "[CLS]...[Person]Stephen Stone[\Person]...[Person]Linda Skyler[\Person]...[SEP]", the relationship between them is "[Relation]Spouse[\Relation]". Formally, after optimization, this sentence is expressed as shown in formula (2):

[0075]

[0076] Meanwhile, in order to analyze the relationship features of the optimized in a more fine-grained manner, the present invention performs syntactic dependency parsing on the original input sentence to extract the dependency path between the head entity and the tail entity. For example, in Figure 2 , the dependency path is The present invention further adopts the average pooling algorithm for all the tokens on the obtained dependency path to generate the dependency feature f dep ∈R d , as shown in formula (3):

[0077]

[0078] where represents the hidden features between the special tokens [E start , [E end , [R start , [R end . In addition, MP(.) and AP(.) represent the max pooling function and the average pooling function, represents the dependency path with length d. For example, in (s i , h i , t i ), the clustering feature f i ∈R d is formed by connecting the features of the entity, the relationship, and the dependencies. This enables f i to consider both semantic and syntactic information simultaneously, as shown in formula (4):

[0079]

[0080] where represents the connection symbol between the respective hidden features.

[0081] Use the BERT language model to obtain the preliminary feature set After that, the purpose of pseudo-relation label clustering in the present invention is to group F into K clustering categories (where K can be manually defined), and these clustering categories can be regarded as the annotations in the training dataset of supervised BERT. Use this training set to fine-tune BERT and then obtain refined labels. According to the experience of some previous clustering works, the present invention adopts an adaptive clustering strategy, including two steps: (1) Use a feature transformation method that transforms from high-dimensional feature representation fi∈Rd to low-dimensional feature representation. Here, a nonlinear mapping network g φ , where is the transformed dimension; (2) Learn a soft assignment to assign all n instances to K centroids to obtain a better clustering effect. This strategy promotes high-confidence assignments and is insensitive to the predefined number of clusters.

[0082] Here, k-means is first applied to obtain K initial centroids Subsequently, using the t-student distribution, the feature similarity between each instance feature and the centroid is calculated here, and then the soft assignment label q i of each sample s ik is calculated, as shown in formula (5):

[0083]

[0084] where α is a hyperparameter used to adjust the clustering algorithm. From another perspective, the present invention also uses the frequency statistics of the entire feature set to be able to make an alternative estimate of the sample distribution p ik , as shown in formula (6):

[0085]

[0086] where is the normalization factor. To adaptively increase the spatial proximity of related entity pairs, use the KL divergence between the soft assignment and the auxiliary distribution to optimize the nonlinear network g φ , and generate a finely optimized pseudo-label l i for each sample s i . Formally, the loss of soft clustering and the pseudo-label l i are as shown in formula (7):

[0087]

[0088] Use these optimized pseudo-labels as self-supervised signals to further fine-tune the small model BERT to obtain a better BERTθ , with relation classification loss As shown in formula (8):

[0089]

[0090] in is the cross entropy loss, and onehot(.) returns a one-hot vector for the pseudo-label. The sentence, head entity, and tail entity are concatenated as input x i , to optimize BERT θ .

[0091] 3) Mentoring Relationship Demonstration Feedback

[0092] like Figure 4 At this stage, we use the fine-tuned BERT θ , in each sample (s i ,h i ,t i ) generates a probability distribution p over all pseudo labels i , where the higher the probability value in a certain relationship category, the clearer BERT’s choice of relationship for this instance is. Then use the entropy H(p i ) is used to measure the uncertainty level p of each instance i , as shown in formula (9):

[0093]

[0094] If the entropy is small, it means high certainty, and if the entropy is large, it means high ambiguity.

[0095] Based on the above considerations, the present invention designs a guiding relationship feedback mechanism in Algorithm 1, which is mainly based on entropy-driven demonstration sample selection (Entropy-Driven Demonstration Selection). Samples with high confidence assignments, whose entropy values ​​are lower than the threshold u, will be clearly placed in the sample demonstration pool A. On the contrary, samples with vague and uncertain relationships need further evaluation. For example, if sample s i Exceeds the entropy threshold u and exhibits a vague relationship r j , then select D relations from pool A corresponding to r j The high probability demonstration of s serves as feedback to the model's first-stage preliminary relationship generation. This process expects the large language model to better adjust s in context learning. i If the number of iterations exceeds the manually set hyperparameter I max , then use formula (7) i To organize the results to produce the final output

[0096] The following table will introduce Algorithm 1 in detail. This algorithm includes an input part, a sample set S, a relationship set L, a probability distribution set P, the number of clusters K, and a demonstration sample set D (lines 1-6); in the first step, a set A is set to represent the total set of relationship guidance examples for all relationship categories (line 7); in the second step, the algorithm traverses the K relationship categories and creates a subset of relationship guidance examples for each relationship category (lines 7-10); in the third step, the algorithm doubly traverses all examples and relationship categories, calculates the entropy of the relationship probability of each example through formula (9), and compares the entropy obtained from the example with a preset threshold. If the relationship that satisfies less than is met, the example and the optimized relationship label are concatenated, and then the concatenated example is added to the subset of the corresponding relationship (lines 11-17); in the fourth step, doubly traverse the number of relationship categories K and the subsets in set A, sort each subset of relationship guidance examples according to the size of the entropy, and finally obtain the demonstration sample set of all relationship categories after sorting (lines 18-20).

[0097]

[0098]

[0099] Example 1

[0100] In the example, an unsupervised open relation extraction text processing method InstructORE proposed by the present invention was used to conduct relevant experiments on two representative datasets, Fewrel and Tacred, and existing unsupervised methods (VAE, Etype+, SelfORE) and weakly supervised methods (RSN, RoCORE, MatchPrompt, ASCORE) were selected as baseline models for comparison to analyze their performance in terms of three metrics: B 3 , V-measure, and ARI. Among them, the B 3 metric is used to evaluate the accuracy of clustering; V-measure is used to measure the balance and integrity of clustering; ARI is used to measure the consistency of clustering.

[0101] The specific implementation is as follows: First, for the experimental sets: (1) Fewrel, which includes 80 relation types and a total of 56,000 instances; (2) Tacred, which includes 41 relation types and 21,773 instances. The experiment divides the Fewrel dataset into 64 predefined relations and 16 new relations, and randomly selects 1,600 instances from the new relations as the test set. For Tacred, the instances marked as no_relation are removed by default, 30 predefined relations and 10 new relations are selected respectively, and 15% of the instances are randomly selected from the new relations as the test set. However, in the experiment, it is not necessary to use the predefined datasets, but all experiments are carried out in an unsupervised manner. The large language model GPT-4 contains rich real-world knowledge and is currently the most excellent generative language model, so the experiment uses GPT-4 as the instruction-based large language model. Set the maximum number of iterations of I max to 8. On the Fewrel and Tacred datasets, the cluster number K is set to 16 and 10 respectively. In Equation (4), the value of α is set to 1. The number of demonstrations D for feedback is set to 6. The maximum entropy u on the relations is set to 2.5. The experiment uses the Stanford text parser to generate dependency trees and uses the 12-layer transformer of the pre-trained bert-base-uncased as the sentence encoder model. To prevent overfitting, all BERT parameters are frozen in the experiment, and only the 8th layer is fine-tuned. AdamW is adopted, and the training batch size and learning rate are 128 and 1e-4 respectively.

[0102] The InstructORE model and the baseline model are tested on the experimental sets. By comparing the performance of the models on three types of metrics, it is found that although no training dataset or manual annotation is used, the performance of InstructORE on all metrics exceeds that of the unsupervised models, and in most metrics, it also exceeds some excellent weakly supervised models such as MatchPromp and ASCORE, and achieves very competitive results in some metrics. Regarding the Fewrel dataset, InstructORE achieves state-of-the-art results in B 3 and V-measure, significantly exceeding ASCORE by 1.3% and 1.8% respectively. Regarding the Tacred dataset, InstructORE in B 3It is superior to all other state-of-the-art models in terms of [aspect], exceeding MatchPrompt by 1.9%, and shows competitive results in terms of the V-measure metric. Then for the ARI metric on the Fewrel and Tacred datasets, although ASCORE and MatchPrompt obtained the highest scores, InstructORE still provided very close results, only 0.6% lower than ASCORE on Fewrel and 1.3% lower than MatchPrompt on Tacred. In addition, in terms of the average score (Avg.), InstructORE showed significant advantages on both datasets, 84.7% and 87.1% on Fewrel and Tacred respectively, both exceeding the highest average score of the current state-of-the-art baseline models.

[0103] Obviously, without additional data or manual annotation, the InstructORE model of the present invention outperforms some current state-of-the-art baseline models, demonstrating strong extraction performance in different metrics and datasets, which proves the effectiveness of the present invention.

[0104] Example 2

[0105] During the continuous interaction between the large language model and the small model, the performance trend of InstructORE was evaluated in this example. For the guided relationship demonstration feedback stage, 2, 4, 6, and 8 demonstration examples were provided in sequence. The observed results were that 3 the F1 scores of [metric] and V-measure continuously increased with the increase of the number of iterations, which was particularly significant in the first few rounds (for example, the increase from the first round to the second round was 5%-8%, and the increase from the second round to the third round was 3%-4%). In addition, an appropriate number of demonstration examples would also significantly affect the context learning ability of the large language model. Compared with InstructORE experiments using 2 or 4 demonstration examples, InstructORE experiments using 6 or 8 demonstration examples obtained an F1 performance 6%-9% higher. In addition, with the increase of the number of iterations, the extraction performance of the models using 6 and 8 demonstration examples tended to converge, indicating the effectiveness of the large language model in low-resource settings.

[0106] Example 3

[0107] To evaluate the generalization ability of InstructORE on different data scales, in this embodiment, 80%, 60%, 40%, and 20% of the data volumes are randomly sampled from Fewrel and Tacred respectively. Since the two baseline models, RoCORE and SelfORE, provide the available latest source codes, InstructORE is compared and evaluated with them here. It is observed that, on the one hand, InstructORE shows strong performance under low-resource settings and stably obtains excellent F1 scores of 76% and 82% with only 20% of the samples. On the other hand, as the sample size decreases, the performance of all three methods will decline. However, compared with RoCORE and SelfORE, InstructORE shows a more gentle decline characteristic. The performance gap between InstructORE and the two baselines expands as the sample size decreases, which also reflects the robustness of the present invention in scenarios with different data volumes.

[0108] Example 4

[0109] To study the discriminative ability of fuzzy relationships with different difficulties, in this embodiment, samples of two different relationships are manually selected, and the effects of three metrics of InstructORE are evaluated under four different difficulty settings of the samples. It is observed that both the first group (Part vs. Subject) and the third group (Sibling vs. Spouse) show low ambiguity, and InstructORE and some of its variant models achieve satisfactory experimental results. On the contrary, the second group (Mother vs. Child) and the fourth group (Headquarters Location vs. Headquarters City) involve semantically similar relationships. InstructORE generally produces relatively excellent results, but in the case of removing the guiding relationship demonstration feedback, the F1 values drop by up to 26.1% and 17.6% at most, which also indicates the importance of the interaction between large language models and small language models.

[0110] Example 5

[0111] To explore the contribution of each semantic feature in the clustering module, in this embodiment, ablation exploration experiments on the semantic features of different clusters are carried out on the two datasets of Fewrel and Tacred, where Rel represents the relationship type feature, Ent represents the entity and entity type feature, Dep represents the dependency path feature between entity pairs, and All represents all features. The observed order of performance decline is as follows:

[0112] InstructORE-All > InstructORE-Rel > InstructORE-Ent > InstructORE-Dep.

[0113] Compared with the InstructORE variant that removes a single other semantic feature, the -Rel variant shows the worst performance, resulting in an average decrease of 8.6% and 6.8% in the three metrics on the Fewrel and Tacred datasets respectively. It can be seen that the relationship types generated by the large language model play a dominant role.

[0114] From the average values of the metrics after removing the entity and entity type features, it can be seen that the performance of InstructORE also decreased significantly by 3.3% and 1.7% without Ent. This fully demonstrates that the entity feature plays a complementary role in the clustering features and the effective supplementation of the entity feature to the overall features.

[0115] When the clustering module for relationship recognition optimization of InstructORE removes the dependency path features, the performance decreases slightly, with the average scores decreasing by 1.3% and 0.8% respectively. This indicates that the dependency path features between entities contribute to a certain extent to solving the problem of missing information on relationship types and entity types.

[0116] Example 6

[0117] In this example, SelfORE and RoCORE are selected as baselines, and three difficult instances are selected to compare with InstructORE. Each instance consists of two sentences with ambiguous relationships. The relationship recognition of these three models for these three difficult instances is specifically analyzed below.

[0118] In Example 1 (across vs. adjacent), both sentences describe the relationship between a river and a location, with subtle semantic differences. Although all models correctly classified the "across" relationship in the first sentence, SelfORE and RoCORE misinterpreted the second sentence. Neither unsupervised self-iteration nor additional transfer learning training data could correct the second sentence from the "across" relationship to the "adjacent" relationship. In Example 2 (part vs. member), SelfORE made completely wrong classifications. It misidentified the "part" relationship in the first sentence as a "member" relationship and the "member" relationship in the second sentence as a "follow" relationship. RoCORE had an incorrect prediction in the second sentence, misidentifying the "member" relationship as a "part" relationship. InstructORE accurately distinguished these two examples by iterating through pseudo-relationship type labels, accurately differentiating the two ambiguous sentences and classifying them into the correct relationship sets. In Example 3 (military rank vs. sound type), although these two relationships were not very ambiguous, the unsupervised method SelfORE made incorrect judgments in both cases, identifying both relationships as "subject" relationships, while the weakly supervised RoCORE corrected these problems. This also reflects that the self-supervised iteration of small language models can have unpredictable errors in some simple instance scenarios, and external supervision is quite useful in solving this problem.

[0119] In summary, the InstructORE model demonstrated higher accuracy, the ability to correct misclassifications, iterative optimization ability, and robustness when dealing with ambiguous relationship instances. These advantages enable InstructORE to perform better in the open relation extraction task.

[0120] Example 7

[0121] This example illustrates the pseudo-relationship optimization process between the entities "D-Day" and "Operation Overlord" in the sentence "After training in the United States, he served in the European theater of operations and received the Distinguished Unit Citation for his actions on D-Day during Operation Overlord" during the iteration of large and small models, where feedback example demonstrations and probabilities for the "part" and "member" categories are given.

[0122] When initially using the large language model GPT-4 to generate relationship and entity pseudo-labels, no relationship guidance demonstration examples are added, but only rely on the design instructions and the capabilities of the large language model itself to generate pseudo-relationships and entity labels, and then officially enter the iteration guided by feedback demonstration examples. In the first round, the probabilities between the relationships "part" and "member" (53% and 37% probabilities) are not very clear yet, and the prediction probabilities of the two only have a 16% gap. The relationship label of the provided demonstration example is "close", which is not very close to the accurate label "part" in terms of semantic similarity and space, but to a certain extent can provide an optimization idea for the representation between the entities "D-Day" and "Operation Overlord" in the follow-up. In the second round, the probabilities between the relationships "part" and "member" (71% and 24% probabilities) expand to 47%. One of the demonstration example relationship labels selected by the small language model BERT is "adjacent", and it gradually approaches the label "part". In the third round, the probabilities between the relationships "part" and "member" (82% and 16% probabilities) further expand to 67%. By the fourth round, the probabilities between the relationships "part" and "member" (82% and 16% probabilities) expand to a very significant 83%, and the example selection of the small language model gives an accurate "part" relationship guidance feedback example after a series of relationship recognition optimizations. Finally, after 5 iterations, the probability distribution becomes very certain that the probabilities of "part" and "member" are 93% and 2% respectively, and the correct result is obtained.

[0123] The above are only the preferred embodiments of the present invention, and do not constitute any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. An unsupervised open relation extraction text processing method based on large and small model interaction, characterized by: The steps include: 1) Preliminary relation generation: Input the sentences and corresponding entities into the large language model to generate preliminary pseudo-relation labels; Using a pre-trained large language model to receive input sentences and entity pairs (h i ,t i ), form instructions c through the designed prompt template i ; These instructions are then used to prompt a large language model to generate preliminary pseudo relation labels and entity pair labels 2) Relationship recognition optimization: Cluster and refine the entity types and relationship labels generated in the first stage, and train a small language model in a self-supervised manner to obtain a fine-grained probability distribution of the output label of each input sentence over all current clustering categories; Use a pre-trained small language model to encode the text description of the entity and obtain a preliminary pseudo label In the original sentence i The context feature f in i ; Cluster the obtained preliminary feature set through an adaptive clustering strategy, and use the clustering category as a training set to fine-tune the small language model to obtain the refined pseudo-label l i ; Use these optimized pseudo-labels as self-supervision signals to further fine-tune the small language model; 3) Guiding relationship demonstration feedback: Select samples with high confidence probability as demonstration examples to provide positive feedback to the large language model for reprocessing instances with uncertain relationships to improve the accuracy and quality of labels; Use the fine-tuned small language model in each sample (s i ,h i ,t i ) of all pseudo labels l i The probability distribution p of the upper generation i Then, based on the entropy-driven demonstration sample selection mechanism, the entropy is used to measure the uncertainty level of each instance and select samples with high confidence probability distribution, that is, the entropy value H(p i ) Samples below the threshold are used as demonstration samples. These samples are placed in the sample demonstration pool to guide the large language model to reprocess instances with uncertain relationships to provide more accurate pseudo-relationship labels.

2. The unsupervised open relation extraction text processing method based on large and small model interaction according to claim 1 is characterized by: Large language models including GPT-4, PaLM 2, and LLaMA 2 are used, and small language models including BERT, MiniLM, and ALBERT are used for interaction.

3. The unsupervised open relation extraction text processing method based on large and small model interaction according to claim 1 or 2, characterized in that: Step 1) Initial relationship generation, as follows: Definition 1 Closed relation extraction: For a given set of sentences Predefined relationship set Extract each sentence s i Medium Entity i and tail entity t i A specific relation r in the relation set R k , where N s Represents the number of sentences in the text collection, N r Represents the number of relations in the relation set, 1≤k≤N r ; Definition 2 Open relation extraction: Open relation extraction is to extract a set of relations from a given text corpus without giving a predefined set of relations R. Head entity h i and tail entity t i The task is formalized as a clustering task, which clusters all the relationship sets corresponding to the current sentence set S. grouped into specific categories; the output of open relation extraction is defined as Where K is the number of categories of defined relationship clusters, and the relationship of each text sentence instance will be placed in one of the cluster categories. For an input sentence s i and the corresponding entity pair h in the sentence i ,t i , design the prompt template, and use s i 、h i and t i Fill the position slots of the template to form instruction c i , and then use these instructions to generate preliminary pseudo relation labels in a large language model and entity pair labels As shown in formula (1): and generate and Entity tag of .

4. The unsupervised open relation extraction text processing method based on large and small model interaction according to claim 1 or 2, characterized in that: Step 2) Relationship identification optimization, as follows: In the prompt command text c i Input the large language model and obtain the preliminary generation result L= Finally, a small language model is used to obtain the contextual features of the preliminary pseudo-labels in the original sentence. Subsequently, the refined soft pseudo-labels are identified through a specific clustering method and further used to fine-tune the small language model to generate the probability of each label. By incorporating special markers [E start ] and [E end ] to update the original input sentence s i , indicating the corresponding head entity and the tail entity The boundary position, in addition, [R start ] and [R end ] is included in the relationship tag The purpose of introducing these special tokens is to obtain richer features of all available information such as text and labels in the small language model. The original sentence s i Combined with the generated relationship labels After the update, the sentence is optimized in form. It is expressed as shown in formula (2): For the optimized The original input sentence is parsed for syntactic dependency to extract the dependency path between the head entity and the tail entity. The average pooling algorithm is then used for all the words on the dependency path to generate the dependency feature f dep ∈R d , as shown in formula (3): in Indicates about special word [E start ]、[E end ]、[R start ]、[R end ], in addition, MP(.) and AP(.) represent the maximum pooling function and the average pooling function, represents a dependency path of length d, in (s i ,h i ,t i ) in which the clustering feature f i ∈R d is formed by connecting the features of entities, relations, and dependencies, which makes f i Both semantic and syntactic information are considered, as shown in formula (4): in Indicates the connection symbols between various hidden features; Use a small language model to obtain a preliminary feature set Finally, the purpose of pseudo-relationship label clustering is to group F into K cluster categories, where K is a fixed value or manually defined, and an adaptive clustering strategy is adopted.

5. The unsupervised open relation extraction text processing method based on large and small model interaction according to claim 4 is characterized by: The adaptive clustering strategy includes two steps: (1) Using high-dimensional features to represent f i ∈R d The feature conversion method to convert to a low-dimensional feature representation uses a nonlinear mapping network g φ ,in is the transformed dimension; (2) Learn a soft assignment to assign all n instances to K centroids, thus achieving better clustering results; this strategy promotes high-confidence assignment and is insensitive to the predefined number of clusters; First, apply k-means to obtain K initial centroids Then, using the t-student distribution, each instance feature is calculated here and centroid The feature similarity of each sample s is then calculated i The soft assigned label q ik , as shown in formula (5): Among them, α is a hyperparameter used to adjust the clustering algorithm, using the frequency statistics of the entire feature set to be able to distribute the sample p ik An alternative estimate is made, as shown in formula (6): in is a normalization factor. In order to adaptively improve the spatial proximity of related entity pairs, soft allocation is used and auxiliary distribution KL divergence between to optimize the nonlinear network g φ , for each sample s i Generate finely optimized pseudo-label l i , formally, the loss of soft clustering is and pseudo-label l i As shown in formula (7): Using these optimized pseudo labels As a self-supervisory signal, further tune the small language model to obtain a better SLM θ , with relation classification loss As shown in formula (8): in is the cross entropy loss, onehot(.) returns a one-hot vector for the pseudo-label, and the sentence, head entity, and tail entity are concatenated as input x i , to optimize SLM θ .

6. The unsupervised open relation extraction text processing method based on large and small model interaction according to claim 5 is characterized by: Step 3) Guidance relationship demonstration feedback, as follows: At this stage, the adjusted SLM is used θ , in each sample (s i ,h i ,t i ) generates a probability distribution p over all pseudo labels i , where the higher the probability value in a certain relationship category, the clearer the small language model is in selecting the relationship for this instance. Then, the entropy H(p i ) is used to measure the uncertainty level p of each instance i , as shown in formula (9): If the entropy is small, it means high certainty, and if the entropy is large, it means high ambiguity.

Citation Information

Patent Citations

  • Open entity relationship extraction method and device, equipment and storage medium

    CN113011189A

  • Rapid relation extraction method based on convolutional neural network and improved cascade labeling

    CN114548090A