A cross-domain small sample relation extraction method and device for learning discriminative semantics and multi-view context
Patent Information
- Application Number
- CN202311564750.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-11-22
AI Technical Summary
在这种场景下,弱化上下文中的关系知识的模型难以区分实体语义相似的不同关系
[0038]2.本方法提出的跨域小样本关系抽取框架包含两个模块。鉴别性的语义对比学习模块放大不同实体之间的语义差异,这可以使模型感知到相似实体的语义之间的细微差异从而具有学习鉴别性实体语义的能力。多视角上下文学习模块以实体信息为基础挖掘实体之间潜在的深层关联关系,并结合整个句子的信息挖掘上下文中的可用于帮助理解关系的信息。这可以进一步利用上下文信息来帮助模型区分具有相似实体语义的不同关系。同时,该模块中的信息过滤机制可以帮助模型在提取上下文中的关系信息时避免过度关注实体信息,从而得到更全面的上下文信息。
Smart Images

Figure CN117852523B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and apparatus for cross-domain few-sample relation extraction by learning discriminative semantics and multi-perspective context, belonging to the field of artificial intelligence natural language processing. Background Technology
[0002] In recent years, with the popularization of the internet and the rapid development of information technology, massive amounts of unstructured text data have emerged online. These include news articles, research publications, blogs, Q&A forums, and social media, generating vast amounts of digital text. Text, as a universal carrier of knowledge, contains a wealth of information within this massive amount of data. Faced with such a vast amount of data, efficiently and accurately extracting the information we want becomes crucial. Relation extraction technology aims to extract the relationships between two specified entities in text. This technology can help us discover connections between things and various concepts, thereby rapidly increasing human knowledge.
[0003] Traditional relation extraction techniques are primarily data-driven, meaning they train relation extraction models on large amounts of labeled training data to learn the meaning of various relations, and then use the trained model to extract relations from text. However, data-driven models heavily rely on large amounts of labeled data, which is often time-consuming and labor-intensive in practice, resulting in insufficient labeled data for model training. This leads to traditional models not learning enough knowledge to complete relation extraction tasks. Therefore, it becomes meaningful to learn knowledge from relation categories in existing labeled data and generalize it to new categories to extract relations of those new types. Thus, few-sample relation extraction techniques are becoming increasingly important. However, in certain specialized fields (medicine, finance, etc.), the available labeled data is even more limited due to the constraints of specialized knowledge. This necessitates learning knowledge from the public domain, where labeled data is readily available, and then generalizing it to new categories specific to the professional domain. Different professional domains have different linguistic features, which means that models trained on the source domain cannot effectively understand relations in the target domain, reducing relation extraction performance. Therefore, how to extract relations in the target domain by training relation instances in the source domain has become an important research direction.
[0004] In recent years, with the development of deep learning technology, various cross-domain few-shot relation extraction methods based on meta-learning frameworks have achieved certain breakthroughs. This allows models to learn cross-task meta-knowledge across multiple relation extraction tasks sampled from large-scale datasets (source domain) and apply it to the target domain. Prototype networks are an efficient algorithm based on meta-learning. They use the class centers of similar samples as relation prototypes and classify the samples to be predicted into the prototypes most similar to them. To improve the quality of relation prototypes, some works learn a large amount of prior knowledge through pre-training, which can help understand the meaning of relations. However, prior knowledge is not directly related to relations, which can only provide limited help for relation extraction. Therefore, some works directly utilize the descriptive information of relations to learn high-quality prototypes.
[0005] Most of these works use an entity-centric matching paradigm to train models to distinguish different relationships, focusing on learning entities rather than context. However, entities in specialized domains often have similar semantic backgrounds. In such scenarios, models that weaken contextual knowledge of relationships struggle to distinguish between different relationships with semantically similar entities. Therefore, existing methods still need improvement to better differentiate between semantically similar relationships. Summary of the Invention
[0006] Existing domain-adaptive few-shot relation extraction methods use an entity-centric matching paradigm to train models to distinguish different relations, focusing on learning entities rather than context. Since the model's learning emphasis is on entity semantics, it determines relations by understanding the head and tail entity classes. When encountering different relations with similar entity semantics, this approach's ability to distinguish between them is limited. Furthermore, this method inherently weakens context learning, ultimately leading to its difficulty in distinguishing between different relations with similar entity semantics.
[0007] To address the aforementioned issues, this invention proposes a cross-domain few-sample relation extraction method based on learning discriminative semantics and multi-perspective context. This method first reduces interference from entities with similar semantics by learning discriminative entity semantics. Furthermore, to enhance the model's ability to learn relational information within the context, the method also performs multi-perspective context learning. It explicitly extracts local and global information from the context to construct multi-perspective contextual representations and uses them as relational information. This effectively utilizes contextual information to distinguish between different relations with similar entity semantics. Simultaneously, the method employs an information filtering mechanism to extract more comprehensive global information for effectively learning relational information within the context.
[0008] This invention proposes a cross-domain few-shot relation extraction method based on learning discriminative semantics and multi-perspective context. This method amplifies the semantic differences between different entities and fully mines relational information within the context based on entity information. The method uses a BERT pre-trained language model as a feature extractor. To accurately extract entity information from sentences, the method uses virtual tokens to emphasize entities and uses the hidden states corresponding to these virtual tokens as entity information. These virtual tokens can utilize undefined tokens reserved in BERT. To learn discriminative entity semantics, the method uses semantic cue templates to extract entity semantics. This guides the model to fully utilize its pre-trained knowledge to represent entities, thereby obtaining entity semantics more accurately and comprehensively. After obtaining entity semantics, this invention amplifies the semantic differences between different entities through semantic contrastive learning. This enables the model to learn discriminative entity semantics, effectively distinguishing different relations with similar entity semantics.
[0009] Furthermore, this method mines global and local information within the context based on entity information to aid in understanding relationships. Local information is obtained by further mining deeper relationships between entities. Global information is mined from the context using head and tail entities to help understand relationships. Global and local information reinforce each other, allowing for further utilization of contextual relationship information to distinguish different relationship categories. Entity information plays a guiding role in mining contextual relationship information, preventing the extraction of irrelevant information. Simultaneously, to avoid over-focusing on entity information and extracting a more comprehensive global information, this method uses an information filtering mechanism to filter out some entity information from the global information. Specifically, the information filtering mechanism calculates the similarity of relationship information in the context obtained based on head and tail entities and uses it as weights. Larger weight values indicate that the obtained relationship information contains more entity information. In this case, a larger proportion of entity information needs to be filtered out to more effectively utilize the extracted relationship information. Finally, this method uses a general training method to obtain a relationship classification loss function to train a model to extract relationships of various categories.
[0010] The technical solution adopted in this invention is as follows:
[0011] A method for cross-domain few-sample relation extraction based on learning discriminative semantics and multi-perspective context includes the following steps:
[0012] Data preprocessing is performed, including appending semantic prompt templates to the end of each sentence in the dataset;
[0013] A cross-domain few-sample relation extraction model is constructed, which includes a feature extraction network, a semantic contrast learning network, a multi-view context learning network, and a relation classification network. The multi-view context learning network includes an information filtering mechanism.
[0014] Using the training set in the dataset, the cross-domain small sample relation extraction model is trained through semantic contrastive learning loss function and relation classification loss function, and the optimal model is obtained using the validation set;
[0015] The optimal model is used to extract relations from sentences in the target domain.
[0016] Furthermore, the data preprocessing also includes: using virtual tags to emphasize the role of entities in relation extraction, and adding virtual tags before and after the head entity and tail entity respectively; the virtual tags use undefined tags reserved in BERT.
[0017] Furthermore, the processing procedure of the semantic contrastive learning network includes: firstly, extracting the semantics of entities using the semantic prompt template, and then amplifying the semantic differences between different entities through semantic contrastive learning, so that the model has the ability to learn discriminative entity semantics and thus effectively distinguish different relationships with similar entity semantics.
[0018] Furthermore, the multi-view context learning network comprises two convolutional neural networks, serving as feature extractors for global and local information, respectively. One convolutional neural network receives entity information and the entire sentence information, extracts contextual information that helps in understanding relationships from the sentence information based on the entity information as global information, and filters out some entity information from the global information through an information filtering mechanism. The amount of filtered information is determined by the similarity weight of the relationship information extracted based on the head and tail entities; the higher the weight, the more entity information is filtered out. The other convolutional neural network only receives head and tail entity information, and obtains potential connections between entities as local information through deep feature extraction. Then, the global and local information are fused to obtain the final contextual relationship information, i.e., multi-view contextual relationship information.
[0019] Furthermore, the relation classification network takes the class center of the relation representation of the same type of relation as the relation prototype, and classifies the relation to be predicted into the relation prototype most similar to it, wherein a parameterless metric function is used to measure the similarity.
[0020] Furthermore, the semantic contrastive learning loss function enables the model to learn discriminative entity representations. Through contrastive learning, it narrows the distance between entities and their corresponding semantics in the representation space, while widening the distance between the semantics of different entities in the representation space. The semantic contrastive learning loss function is expressed as follows:
[0021]
[0022] in, and These represent the head and tail entity information of the k-th sample in the i-th relation category, respectively. and Let represent the corresponding entity semantic information, and d be the standardized cosine similarity distance function. This means that the average of the contrastive learning losses obtained from k sets of data is taken as the final semantic contrastive loss.
[0023] Furthermore, the relation classification loss function enables the model to extract relations of different categories. In a task consisting of a sampled set of data, samples in the query set are classified into a certain category prototype in the support set, and the model is trained on multiple sets of repeatedly sampled tasks. The relation classification loss function is expressed as follows:
[0024]
[0025] Where, q j For the j-th sample in the query set, p q This indicates the relational prototype category to which the sample belongs, τ is a temperature parameter used to control the smoothness of similarity, d represents the standardized cosine similarity distance function, and G represents the total number of samples to be classified.
[0026] Furthermore, the weight calculation method in the information filtering mechanism is as follows:
[0027]
[0028]
[0029]
[0030] in, and These represent the contextual relationship information extracted from the head and tail entities using a convolutional neural network, respectively. w is the similarity weight, and HGI and TGI are the relationship information after information filtering. The two are combined to obtain the final global information GI.
[0031] A cross-domain few-sample relation extraction device that learns discriminative semantics and multi-perspective context, comprising:
[0032] The data preprocessing module is used to perform data preprocessing, including concatenating semantic prompt templates to the end of each sentence in the dataset;
[0033] The model building module is used to build a cross-domain few-sample relationship extraction model. The cross-domain few-sample relationship extraction model includes a feature extraction network, a semantic contrast learning network, a multi-view context learning network, and a relationship classification network. The multi-view context learning network includes an information filtering mechanism.
[0034] The model training module is used to train the cross-domain small sample relation extraction model using the training set in the dataset, through the semantic contrast learning loss function and the relation classification loss function, and to obtain the optimal model using the validation set.
[0035] The relation extraction module is used to extract relations from sentences in the target domain using the optimal model.
[0036] Key aspects of this invention include:
[0037] 1. This method proposes discriminative entity semantic learning and multi-perspective context learning, which integrates the learned entity information and context information into a single model. This can effectively avoid the model confusing different relationships with similar entity semantics, and enable the model to more effectively generalize from general domains to specific professional domains.
[0038] 2. The proposed cross-domain few-sample relation extraction framework comprises two modules. The discriminative semantic contrast learning module amplifies the semantic differences between different entities, enabling the model to perceive subtle semantic differences between similar entities and thus learn discriminative entity semantics. The multi-view context learning module mines potential deep relationships between entities based on entity information and combines this with information from the entire sentence to extract contextual information that can aid in understanding the relationships. This further utilizes contextual information to help the model distinguish different relationships with similar entity semantics. Simultaneously, the information filtering mechanism in this module helps the model avoid overemphasizing entity information when extracting relational information from the context, thereby obtaining more comprehensive contextual information.
[0039] 3. This method proposes two interacting and mutually influential loss functions: semantic contrastive learning loss function and relation classification loss function.
[0040] Compared with the prior art, the positive effects of the present invention are as follows:
[0041] 1. This invention addresses the issue that current cross-domain small sample relation extraction methods are affected by the high semantic similarity of entities in professional domains, resulting in the lack of discriminative entity semantics. It proposes a discriminative entity semantic learning module, which uses semantic contrastive learning to enable the model to learn discriminative entity representations, thereby improving the model's ability to identify different relations with similar entity semantics.
[0042] 2. In response to the problem that current methods for learning relational information in sentences cannot fully learn relational information in the context using an entity-centered learning approach, this invention proposes a multi-perspective contextual learning module. This module extracts local relational information based on head and tail entities and extracts global relational information based on sentence information of entities. The two modules complement each other and combine to form multi-perspective contextual information, thereby improving the model's ability to distinguish different relations using contextual information.
[0043] 3. This invention provides a weighted adaptive information filtering mechanism that can filter out some entity information based on the similarity of the extracted context information, thereby avoiding excessive focus on entities during the learning process and obtaining more comprehensive contextual knowledge. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the method of the present invention;
[0045] Figure 2 This is a schematic diagram of the framework structure proposed by the method of this invention. Detailed Implementation
[0046] To better illustrate the cross-domain few-sample relation extraction method based on learning discriminative semantics and multi-perspective context proposed in this invention, the invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0047] Figure 1 This is an overall flowchart of a cross-domain few-sample relation extraction method based on learning discriminative semantics and multi-perspective context according to the present invention, which includes four parts: data preprocessing, initializing the model framework, model training, and relation extraction.
[0048] Step 1. Data Preprocessing. The manually designed entity semantic prompt template is appended to the end of each sentence in the dataset.
[0049] Step 2. Initialize the model framework. Figure 2 This is the model framework designed in this invention, which includes a semantic contrast learning module, a multi-view context learning module, an information filtering mechanism, and a relationship classification module.
[0050] Step 3. Model Training. This invention trains the model using semantic contrastive loss and relational classification loss. On the validation set, when the total loss function converges and the model achieves optimal performance on the validation set, its parameters are saved as the optimal model. During validation and testing, this method only inputs data into a multi-view context module to extract multi-view contextual information.
[0051] Step 4. Relation Extraction. Using the optimal model obtained in Step 3, the test set data is used as input to the model. The relation in each sentence is obtained by combining multi-perspective contextual information and entity information. Then, the similarity between the relation to be classified and the given relation category is calculated using a metric function, and the category is assigned to the relation most similar to it.
[0052] According to the solution provided by the present invention, the specific steps of a method for cross-domain few-sample relation extraction based on learning discriminative semantics and multi-perspective context, according to an embodiment of the present invention, are as follows:
[0053] Step 1. Data Preprocessing. Since this method requires semantic prompt templates to guide the model in providing the semantics of entities, this invention designs semantic prompt templates and adds them to each data entry. Secondly, this invention uses undefined tags within BERT as virtual tags to emphasize the role of entities in relation extraction, and adds them before and after the head and tail entities, respectively.
[0054] Step 2. Initialize the model framework. This method uses metric learning as the basic model framework, a popular meta-learning-based framework. It uses the BERT pre-trained language model as the feature encoder. The model input is mainly divided into two parts: a support set and a query set. The ultimate goal of the model is to classify the query set into a specific relation category within the support set.
[0055] The method includes:
[0056] First, the sampled support set and query set data are input into the feature encoder to obtain feature codes. Due to the addition of virtual entity markers and semantic cue templates, the feature codes obtained by this method are both entity feature codes and semantic feature codes.
[0057] Then, semantic contrastive learning amplifies the semantic differences between different entities. This process is accomplished using the loss function (semantic contrastive learning loss) designed in this invention. Moreover, this process is only performed on the support set because the class relationships between data in the query set cannot be utilized.
[0058] For the part learning multi-view context, two convolutional neural networks (CNNs) are first initialized as feature extractors for global and local information. Each CNN has four layers with 256, 128, 64, and 1 kernels respectively, and the kernel size is set to 5. One CNN receives entity information and the entire sentence information, extracting contextual information that helps understand relationships as global information based on the entity information. During this process, all data in the support set and query set undergoes information extraction. The extraction of global information may be significantly influenced by entity information, resulting in a large amount of entity information included in the extracted global information. This reduces the effectiveness of relational information in the context. Therefore, this method uses an information filtering mechanism to filter out some entity information from the global information. This helps to obtain more comprehensive global information. The amount of filtered information is determined by the similarity weights of the relational information extracted based on the head and tail entities. Higher weights result in more entity information being filtered out. The other CNN only receives head and tail entity information, extracting potential connections between entities through deep feature extraction as local information. Then, this method fuses the global and local information to obtain the final relational information in the context, i.e., multi-view contextual relational information. Then, it is combined with entity information to obtain a relational representation of a sentence.
[0059] Finally, relationship classification is performed. This method uses the class center of the relationship representation of similar relationships as the prototype of that type of relationship, and classifies the relationship to be predicted into the most similar relationship prototype. For similarity measurement, a parameter-free metric function, such as cosine similarity, is used. This process is accomplished through a classification loss function.
[0060] Step 3. Model Training. This method trains the model by optimizing two loss functions.
[0061] The first loss function, the semantic contrastive learning loss function, enables the model to learn discriminative entity representations. Specifically, this loss function is obtained from the contrastive learning framework. This method uses contrastive learning to bring entities and their corresponding semantics closer together in the representation space, and to widen the distance between the semantics of different entities in the representation space. This process amplifies the semantic differences between different entities, thereby learning discriminative entity representations. Contrastive learning is accomplished through data grouping. Each group of contrastive learning data consists of one data point from each relation category in the support set. This invention uses the average of the loss functions obtained from each group of contrastive learning as the final contrastive learning loss function. Contrastive learning requires calculating the distance, or similarity, between entities and their semantics, as well as between the semantics of different entities, in the feature space. The distance is calculated using a parameterless distance function, such as cosine similarity.
[0062] The second loss function, the relation classification loss function, enables the model to extract relations of different categories. Specifically, this loss function is obtained from the meta-learning framework. That is, in a task consisting of a sampled set of data, the model needs to classify samples in the query set into a certain category prototype in the support set. The model is trained on multiple sets of repeatedly sampled tasks. Samples in the support set need to first learn the prototype of each relation category before participating in the classification process. The relation prototype can be considered a standard for this category of relation. During the classification process, a parameterless distance function is used to calculate the similarity between relations, and a temperature parameter is used to control the smoothness of the probability. The loss function obtained from this process represents the model's classification ability. Optimizing the loss function maximizes the probability that a sample to be classified is classified into the correct relation category and minimizes the probability of being classified into other relation categories.
[0063] This method combines two loss functions into a total model loss function by assigning different weights. The invention is implemented on a validation set. When the total loss function converges and achieves optimal performance on the validation set, the model parameters are saved as the optimal model. During validation and testing, this method only inputs data into a multi-view context module to extract multi-view contextual information.
[0064] Step 4. Relation Extraction. Using the optimal model obtained in Step 3, the test set data is used as input to the model. The relation in each sentence is obtained by combining multi-perspective contextual information and entity information. Then, the similarity between the relation to be classified and the given relation category is calculated using a metric function, and the category is assigned to the relation most similar to it.
[0065] For example, in step 1, a sentence with added semantic cue templates is: "[E1]Newton[\E1]served as the president of[E2]the Royal Society[\E2].Newton means[M1],the Royal Society means[M2]". The first half of the sentence is a sentence from a dataset, where [E1], [\E1], [E2], and [\E2] represent undefined dummy tags used to emphasize the role of the entity in the relation. The second half of the sentence is a semantic cue template, where [M1] and [M2] represent the semantics of the head entity and the tail entity, respectively.
[0066] For example, in step 2, the convolutional neural network consists of multiple convolutional layers, activation layers, and dropout layers. Each convolutional layer extracts features from the input through a convolution operation between the convolutional kernel and the input. The activation layer uses an activation function to increase the non-linear expressive power of the model, and the dropout layer can alleviate the overfitting problem caused by insufficient samples to some extent.
[0067] For example, in step 3, the semantic contrastive learning loss function is expressed as:
[0068]
[0069] in, and HS represents the head entity and tail entity information of the k-th sample in the i-th relation category, respectively. k i and TS k i Let represent the corresponding entity semantic information, and d be the standardized cosine similarity distance function. This means that the average of the contrastive learning losses obtained from K sets of data is taken as the final semantic contrastive loss, where N represents the number of relation categories.
[0070] For example, in step 3, the weight calculation method in the information filtering mechanism is as follows:
[0071]
[0072]
[0073]
[0074] in, and These represent the contextual relationship information extracted using a convolutional neural network based on the head and tail entities, respectively. `w` represents the similarity weight. `HGI` and `TGI` represent the relationship information after information filtering. Combining these two yields the final global information, `GI`.
[0075] For example, in step 3, the final representation of the relations implied in each sentence is as follows:
[0076]
[0077]
[0078] Where R represents the relational information contained in the sentence, Head and Tail represent the head entity and tail entity respectively, and MC is the multi-perspective contextual information, which is obtained from the local information LI and the global information GI.
[0079] For example, in step 3, the relationship classification loss function is expressed as follows:
[0080]
[0081] Where, q j For the j-th sample in the query set, p qThis indicates the relational prototype category to which the sample belongs, τ is a temperature parameter used to control the smoothness of similarity, d represents the standardized cosine similarity distance function, and G represents the total number of samples to be classified.
[0082] For example, in step 3, the total loss function of the model is expressed as follows:
[0083] L=λ1L E +λ2L FS
[0084] Wherein, λ1 and λ2 are the weights of the semantic contrastive learning loss and the relation classification loss, respectively. These two weights are manually determined based on the model's performance on the validation set.
[0085] Table 1 shows the comparative experimental data of the present invention and existing methods.
[0086] Table 1
[0087]
[0088] The experimental results above demonstrate that the present invention (MCDS) outperforms existing methods in various task scenarios. Its advantages are particularly pronounced in tasks where there is only one sample per relation category.
[0089] Another embodiment of the present invention provides a cross-domain few-sample relation extraction device for learning discriminative semantics and multi-perspective context, comprising:
[0090] The data preprocessing module is used to perform data preprocessing, including concatenating semantic prompt templates to the end of each sentence in the dataset;
[0091] The model building module is used to build a cross-domain few-sample relationship extraction model. The cross-domain few-sample relationship extraction model includes a feature extraction network, a semantic contrast learning network, a multi-view context learning network, and a relationship classification network. The multi-view context learning network includes an information filtering mechanism.
[0092] The model training module is used to train the cross-domain small sample relation extraction model using the training set in the dataset, through the semantic contrast learning loss function and the relation classification loss function, and to obtain the optimal model using the validation set.
[0093] The relation extraction module is used to extract relations from sentences in the target domain using the optimal model.
[0094] For the specific implementation process of each module, please refer to the description of the method of the present invention above.
[0095] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.
[0096] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) storing a computer program that, when executed by a computer, implements the various steps of the method of the present invention.
[0097] Although the specific details, implementation algorithms, and accompanying drawings of the present invention have been disclosed for illustrative purposes to aid in understanding and implementing the invention, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. The invention should not be limited to the content disclosed in the preferred embodiments and accompanying drawings; the scope of protection claimed by the invention is defined by the claims.
Claims
1. A method for cross-domain few-sample relation extraction based on learning discriminative semantics and multi-perspective context, characterized in that, Includes the following steps: Data preprocessing is performed, including appending semantic prompt templates to the end of each sentence in the dataset; A cross-domain few-sample relation extraction model is constructed, which includes a feature extraction network, a semantic contrast learning network, a multi-view context learning network, and a relation classification network. The multi-view context learning network includes an information filtering mechanism. Using the training set in the dataset, the cross-domain small sample relation extraction model is trained through semantic contrastive learning loss function and relation classification loss function, and the optimal model is obtained using the validation set; The optimal model is used to extract relations from sentences in the target domain; The processing of the semantic contrastive learning network includes: firstly, extracting the semantics of entities using the semantic prompt template, and then amplifying the semantic differences between different entities through semantic contrastive learning, so that the model has the ability to learn discriminative entity semantics and thus distinguish different relationships with similar entity semantics; The multi-view context learning network comprises two convolutional neural networks (CNNs), serving as feature extractors for global and local information, respectively. One CNN receives entity information and the entire sentence information, extracting contextual information that helps in understanding relationships from the sentence information based on the entity information as global information. An information filtering mechanism then filters out some entity information from the global information; the amount of filtered information is determined by the similarity weights of the relationship information extracted from the head and tail entities—higher weights result in more filtered-out entity information. The other CNN only receives head and tail entity information, extracting potential connections between entities through deep feature extraction as local information. Finally, the global and local information are fused to obtain the final contextual relationship information, i.e., the multi-view contextual relationship information. The weight calculation method in the information filtering mechanism is as follows: in, and These represent the contextual relationship information extracted from the head and tail entities using a convolutional neural network, respectively. For similarity weights, and The relationship information is filtered and then combined to obtain the final global information GI.
2. The method according to claim 1, characterized in that, The data preprocessing also includes: using virtual tags to emphasize the role of entities in relation extraction, and adding virtual tags before and after the head entity and tail entity respectively; the virtual tags use undefined tags reserved in BERT.
3. The method according to claim 1, characterized in that, The relation classification network uses the class center of the relation representation of the same type of relation as the relation prototype, and classifies the relation to be predicted into the relation prototype most similar to it, wherein a parameterless metric function is used to measure the similarity.
4. The method according to claim 1, characterized in that, The semantic contrastive learning loss function enables the model to learn discriminative entity representations. Through contrastive learning, it narrows the distance between entities and their corresponding semantics in the representation space, and widens the distance between the semantics of different entities in the representation space. The semantic contrastive learning loss function is expressed as follows: in, and These represent the head and tail entity information of the k-th sample in the i-th relation category, respectively. and For the corresponding entity semantic information, For the standardized cosine similarity distance function, Indicates will The mean of the contrastive learning loss obtained from the sets of data is taken as the final semantic contrastive loss; The relation classification loss function enables the model to extract relations of different categories. In a task consisting of a sampled set of data, the samples in the query set are classified into a certain category prototype in the support set, and the model is trained on multiple sets of repeatedly sampled tasks. The relation classification loss function is expressed as follows: in, For the j-th sample in the query set, This indicates the relation prototype category to which the sample belongs. Temperature is a parameter used to control the smoothness of similarity. d Represents the standardized cosine similarity distance function. G This represents the total number of samples to be classified.
5. A cross-domain few-sample relation extraction device for learning discriminative semantics and multi-perspective context, characterized in that, The apparatus for performing the method according to any one of claims 1 to 4 comprises: The data preprocessing module is used to perform data preprocessing, including concatenating semantic prompt templates to the end of each sentence in the dataset; The model building module is used to build a cross-domain few-sample relationship extraction model. The cross-domain few-sample relationship extraction model includes a feature extraction network, a semantic contrast learning network, a multi-view context learning network, and a relationship classification network. The multi-view context learning network includes an information filtering mechanism. The model training module is used to train the cross-domain small sample relation extraction model using the training set in the dataset, through the semantic contrast learning loss function and the relation classification loss function, and to obtain the optimal model using the validation set. The relation extraction module is used to extract relations from sentences in the target domain using the optimal model.
6. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Text information extraction model training method, text information extraction method and application
CN115270801A
Method for extracting entity relationship based on context dependency perception graph convolutional network
CN116992881A