Small sample unified granularity relation extraction method based on large language model
By introducing contextual learning and thought chain technology, the fragmentation problem of relationship extraction algorithms in different text ranges is solved, the accuracy of relationship extraction of large language models in small sample scenarios is improved, and the unification of sentence-level, document-level and cross-document-level relationship extraction is achieved.
Patent Information
- Application Number
- CN202510091857.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-09-12
AI Technical Summary
Existing relation extraction algorithms are fragmented in sentence-level, document-level, and cross-document-level tasks, and datasets are scarce in small sample scenarios, resulting in poor application of models in different text ranges, especially the direct application of large language models in relation extraction tasks.
By introducing contextual learning, thought chain and entity enhancement technologies, and transforming the relationship extraction task into a sequence generation task, the powerful reasoning ability of the large language model is utilized to uniformly solve the sentence-level, document-level and cross-document-level relationship extraction problems, including task information description, context example demonstration, context entity enhancement and thought chain output.
It improves the accuracy and effectiveness of relationship extraction of large language models in small sample scenarios. Through context learning and thought chain technology, it enhances the model's reasoning ability at different granularities and achieves the effectiveness of unified granularity relationship extraction.
Smart Images

Figure CN120633654A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence and natural language processing, and specifically relates to a small sample unified granularity relationship extraction method based on a large language model. Background Art
[0002] The relationship extraction task is usually a subsequent step of named entity recognition, that is, given an entity and its type, the possible relationships between them are extracted. In a broad sense, relationship extraction does not restrict the scope of text within which the entities that constitute a certain relationship exist. However, most traditional relationship extraction algorithms restrict it to within sentences and take this restriction as the default condition. Recent work has gradually recognized the huge limitations of this restriction, and thus derived document-level relationship extraction and cross-document-level relationship extraction compared to sentence-level relationship extraction. However, these studies are extremely fragmented. All algorithms proposed for document-level and cross-document-level relationship extraction tasks are domain-specific, that is, these algorithms cannot be applied to relationship extraction tasks in different text ranges without modification.
[0003] Sentence-level, document-level, and cross-document-level relationship extraction tasks are essentially all broadly defined as relationship extraction tasks. A more practical algorithmic model should be able to accurately extract the relationships within the text, regardless of the context in which the entities are located. Furthermore, for both sentence-level, document-level, and cross-document-level relationship extraction tasks, currently available public datasets are scarce. Labeling relationship extraction datasets in vertical fields requires extensive manual reasoning, which is time-consuming and labor-intensive. Therefore, the algorithmic model must be effective in small sample scenarios.
[0004] With the emergence of ChatGPT, the massive number of parameters in stacking a large number of Transformer Decoders and unsupervised training on ultra-large data sets have transformed language models from quantitative to qualitative, resulting in the so-called Large Language Model (LLM). LLMs have demonstrated excellent performance in various downstream tasks, especially in small sample sizes. However, the straightforward application of LLMs to the downstream task of relation extraction has yielded poor results, requiring corresponding hint learning techniques to improve accuracy. Summary of the Invention
[0005] Purpose of the invention: The present invention addresses the fragmented research issues of sentence-level relation extraction, document-level relation extraction, and cross-document-level relation extraction, introduces contextual learning, thought chaining, and entity enhancement technologies, and gives full play to the powerful reasoning ability of the large language model. By converting the relation extraction task into a sequence generation task, the sentence-level, document-level, and cross-document-level relation extraction problems are uniformly solved with one model; in addition, the introduction of contextual learning, thought chaining, and entity enhancement technologies improves the accuracy of the large language model in relation extraction tasks and its effectiveness in small sample scenarios.
[0006] Technical Solution: To achieve the above-mentioned purpose, the present invention proposes a method for extracting small sample uniform granularity relationships based on a large language model, comprising the following steps:
[0007] (1) Task information description: provide a specific task description as part of the input context of the large language model;
[0008] (2) Contextual example demonstration: Give the large language model a comparable example as a contextual demonstration and incorporate it into the input in the form of linear natural language;
[0009] (3) Contextual entity enhancement: Entity enhancement is performed on the context input to the large language model, thereby better indicating the location information of the large language model entity in the context;
[0010] (4) Thinking chain output defines a way to serialize relation triples and incorporates additional reasoning information that helps with relation extraction, thereby helping the large language model complete relation extraction in steps.
[0011] Furthermore, in step (1), a unified relationship extraction task definition is given that can encompass three relationship extraction tasks at the sentence level, document level, and cross-document level.
[0012] Furthermore, in step (2), the context demonstration example is generated by a randomly selected sample from a training set of a given data set through a prompt template constructor.
[0013] Furthermore, in step (3), entity enhancement only adds an identifier at the end of the entity, rather than adding some identifier symbol at both the head and the end of the entity.
[0014] Furthermore, in step (4), the infix form of the relation triple is converted to a suffix form. First, the first mention of the head entity is concatenated with the type output, and then the sentence number in which the entity appears in the context is added. The sentence number starts from 0, which helps the large language model focus on the sentences related to the entity, thereby improving the reasoning ability of relation extraction; followed by a similar tail entity representation; and finally, the relation representation surrounded by “@” is output.
[0015] Beneficial Effects: This invention effectively improves the effectiveness of uniform-granularity relation extraction tasks in small sample scenarios. First, it introduces contextual learning technology, which not only helps large language models better understand tasks by providing specific task descriptions, but also allows them to learn given tasks from a few demonstration examples. Second, it introduces thought chain technology, which incorporates additional information that aids relation extraction into relation triples, thereby helping large language models complete relation extraction in steps and further enhancing their reasoning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is the overall framework diagram of the method of the present invention;
[0017] Figure 2 For regular expressions Visualization diagram;
[0018] Figure 3 For regular expressions Visualization diagram;
[0019] Figure 4 For regular expressions Visualization diagram;
[0020] Figure 5 This is a graph of the step learning rate adjustment strategy in the method of the present invention. DETAILED DESCRIPTION
[0021] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0022] This paper proposes a small-sample unified granularity relationship extraction model based on a large language model. For the unified granularity relationship extraction task, four basic concepts of the task are introduced: context, entity, mention, target entity, and relation.
[0023] Context: Context Refers to a continuous linear sequence of words, w i Represents the i-th word in the context, n represents the length of the context, for sentence-level relation extraction, It consists of a sentence with an average number of words in the order of tens; for document-level relation extraction, It consists of the order of words in a document, and the average number of words is in the hundreds. For cross-document level relationship extraction, It consists of the order of words in documents of several document paths, with an average number of words in the order of thousands.
[0024] Mention: Mention is determined by context. The subsequences in the word sequence are represented as m={w i ,…,w i+k-1}. Which means that the context where k is the starting index in , and k is the length of the word subsequence corresponding to the mention.
[0025] Entity: An entity is an abstract concept consisting of all mentions that semantically refer to the same identifier. Therefore, entity e = {m|m has the same semantics} is a non-empty set of mentions. An entity has at least one mention. Each entity has its type. The function et(e) represents the type of entity e. et(e)∈{et1,…,et m} is an element in the predefined entity type set, and m represents the number of elements in the entity type set.
[0026] Target Entity: In the context The entities that appear in are not necessarily the entities that need to be extracted, but may only be in the context Therefore, the target entity t refers to the entity whose possible relationship with other entities needs to be extracted.
[0027] Relation: A relation is a triple consisting of an entity and a predefined relationship. <e h ,r,e t >, indicating the head entity e h With the tail entity e t There is a relationship r between them, where h represents the head and t represents the tail. It is an element in the predefined relationship set, and K represents the number of elements in the relationship set.
[0028] Based on the basic concepts defined above, the specific definition of the unified relation extraction task is given below.
[0029] Unified Relation Extraction: Given Context Entity Collection where n e Indicates the number of entities, target entity set where n tRepresents the number of target entities, and the goal is to extract all possible relationship triples in ε
[0030] This method uses the powerful reasoning ability of large language models to improve the effectiveness of relation extraction tasks at uniform granularity in small sample scenarios. Figure 1 As shown, the complete process of the present invention includes task information description, context example demonstration, context entity enhancement, and thought chain output. The specific implementation method is described as follows:
[0031] The task information description stage corresponds to step (1) of the technical solution. The specific implementation method is as follows: first, a brief description of the unified granularity relationship extraction task is given, and the task description instruction I is defined as a natural language sequence, which is specifically defined as follows:
[0032] I=Given a context composed of a sequence of words and the entities
[0033] withinit,predict pairs of entities with relationships and therelations.
[0034] I = Given a context consisting of a sequence of words and the entities present in it,
[0035] Predict the relationship triples (head entity, tail entity, relationship)
[0036] The unified granularity relation extraction instruction is "given a context consisting of a word sequence and the entities therein, predict the relation triples (head entity, tail entity, relation) therein."
[0037] The context example demonstration phase corresponds to step (2) of the technical solution. The specific implementation plan is: a sample randomly selected from the training set of a given dataset is used to generate a context demonstration example D through the prompt template constructor, which is specifically defined as follows:
[0038]
[0039] where f prompt represents the prompt template constructor to be introduced in step (3) of the technical solution, ε represents the target entity provided in the sample, and f enhanced represents the entity enhancement function to be introduced in step (3) of the technical solution, and Represent the context and entity of the sample respectively, Represents the string concatenation operator, f relationrepresents the linearization function of the relationship introduced in step (4) of the technical solution, Represents the set of relation triplets of sample annotations.
[0040] Therefore, the context demonstration example D is formed by concatenating the sample examples after prompt template transformation and the linearized relation triples.
[0041] The context entity enhancement phase corresponds to step (3) of the technical solution. The specific implementation scheme is: given an entity mention m = {w i ,…,w i+k-1} and the entity type et, define the function as follows:
[0042]
[0043] The function ψ takes as input a sequence of words mentioning m and the type of the entity it belongs to, et, and outputs an augmented natural language sequence. Specifically, if the dataset provides entity type labels, the augmented natural language sequence is constructed by appending the type information, "@et@," surrounded by @ symbols. Otherwise, the augmented natural language sequence is constructed by appending only the @ symbols.
[0044] Therefore, the function of entity enhancement of context using ψ function The definition of is as follows: First, traverse the entity set in the sample Record the location and type information of each mention in the dictionary E s Then, traverse the original context For each word in , if the word is the actual word of a mention, then the mention is marked using the ψ function and added to the end of the new context, otherwise it is directly added to the new context. Finally, the entity-enhanced context is returned
[0045] After obtaining the task description instruction I, context demonstration example D and entity enhancement context After that, define the prompt template conversion function f prompt as follows:
[0046]
[0047] Entities:, Context:, and Relations: are all prompts inserted into the prefix. entity (ε) represents the linearization function of the target entity set, Indicates string concatenation operation. entity The function separates all word sequences mentioned in the entity with the symbol "&" and finally concatenates the entity type surrounded by the "@" symbol as the end of the sequence.
[0048] Finally, the task description instruction I, the context demonstration example D, and the prompt template are concatenated to obtain the input of the large language model.
[0049] The output stage of the thinking chain corresponds to the technical solution step (4). The specific implementation plan is: define the function The relation triple <e h ,r,e t Linear serialization into a natural language string, which is defined as follows:
[0050] δ(e h )=ψ(m h ,e h )s h
[0051] δ(e h ) represents the entity serialization representation of the thought chain information, where ψ is the function defined in step (3) of the technical solution, s h Represents entity e h The set of sentences in which the entity appears. The delta function lists all mentions of the entity, followed by the entity type, and finally the sentence number in which the entity appears in the context.
[0052]
[0053] function Convert the infix form of the relation triple to a suffix form. First, concatenate the first mention of the head entity with the type output, then add the sentence number (starting from 0) in which the entity appears in the context. This helps the large language model focus on sentences related to the entity, thereby improving the reasoning ability of relation extraction. This is followed by a similar tail entity representation. Finally, the relation representation surrounded by "@" is output.
[0054] Utilization function All the relation triples contained in the context Serialized into the output string Y of the large language model, specifically defined as follows:
[0055]
[0056] The function "x".join(y) means inserting the string "x" between the elements y of the set and concatenating them into a string. From the above formula, we can see that the output Y is composed of the relation triple set Each triple in is serialized and separated by spaces and concatenated together to form a text string.
[0057] In order to convert the generated serialized natural language string into structured data that can be automatically evaluated for accuracy, the string needs to be deserialized into relation triples.
[0058] Defining regular expressions Represents a character sequence surrounded by the @ symbol, which can match the type string at the end of the entity or the generated relationship type string. The specific definition is given by the following formula. Figure 2 Regular Expression The matching pattern is visualized.
[0059]
[0060] Defining regular expressions Represents the relation triple string after matching serialization Regular expression for . First, match any character sequence ending with the "#" character, which represents the head and tail entities e h and e t Serialized string; then match any whitespace; finally match the character sequence surrounded by the @ symbol (ie ), which represents a relationship type string. The specific definition of is given by the following formula. Figure 3 Regular Expression The matching pattern is visualized, where group#alpha represents
[0061]
[0062] Defining regular expressions Indicates the entity sequence ψ(m,e t ) is the regular expression of the string after the thought chain #sent(e)#. First, match any character sequence surrounded by whitespace, indicating an entity mention string; then match a character sequence surrounded by the @ symbol (i.e. ), indicating the type information of the entity; and finally matching the character sequence surrounded by the "#" symbol, indicating the thought chain information. The specific definition of is given by the following formula. Figure 4 Regular Expression The matching pattern is visualized, where group#alpha represents
[0063]
[0064] The deserialization algorithm first passes For input string Perform pattern matching to obtain the head and tail entity strings s in the serialized strings of all relationship triples r and relationship type r; then by For the head and tail entity string s r Perform pattern matching to obtain the mention string s corresponding to the head and tail entities m 、Type e t , thought chain sentences sents, and add all mentions to the set m s In the above example, we get the mention set e of the head and tail entities. s Finally, any two mentions of the head and tail entities are combined into a relation triple <m h ,r,m t >Add to output collection
[0065] This paper proposes a small-sample unified granularity relationship extraction method based on a large language model. To test the accuracy and versatility of this method, experiments were conducted on sentence-level, document-level, and cross-document-level relationship extraction datasets. The sentence-level relationship extraction task includes the SemEval-2010 and NYT datasets; the document-level relationship extraction task dataset is Re-DocRED; and the cross-document-level relationship extraction task dataset is CodRED. This method was evaluated using the Precision, Recall, and F1 metrics and compared with other relationship extraction methods. In addition, to verify the effectiveness of this method in small-sample scenarios, 100, 200, 10%, 20%, and 30% samples were randomly extracted from the full Re-DocRED dataset as small-sample datasets for experiments. The higher the value of the above metrics, the better the experimental results.
[0066] The large language models used are Meta's open-source LLaMA-2 versions 7b and 13b, which have 7 billion and 13 billion parameters, respectively, and 32-layer and 40-layer Transformer Decoders, respectively.
[0067] When training the model, we use Figure 2 Assume that the initial learning rate is 4e-5, and the hyperparameters of the strategy are adjusted so that the learning rate is reduced by 0.8 times after every 100 epochs, and the total number of epochs is 1000. Then the learning rate change trend is as follows: Figure 5 shown.
[0068] All model hyperparameter settings are shown in Table 1. The model uses AdamW as the parameter optimizer; the initial learning rate of the model parameters is set to 5e-5; the learning rate adjustment strategy is StepLR, with a step size of 1 and a gamma of 0.8; the model hidden layer vector dimensions are 4096 and 5120; the maximum length of the model input context is set to 4096; the number of samples per batch is set to 4 when training the sentence-level relation extraction task, and the number of samples per batch is set to 1 when training the document-level and cross-document-level relation extraction tasks; and the number of training epochs is set to 10.
[0069] The model was trained using the Parameter Efficient Fine Tuning (PEFT) strategy and the LoRA (Low-Rank Adaptation) technique. During training, the parameters of the LLaMA itself were frozen, and only the low-rank transformation matrix of its sub-model was trained. The parameters were set to r = 512, alpha = 1024, the proxy sub-models were "q_proj" and "v_proj", and dropout was set to 0.05.
[0070] Table 1 UniGReM hyperparameter setting information
[0071]
[0072]
[0073] The comparative experimental results of different relation extraction methods on sentence-level, document-level, and cross-document-level relation extraction datasets are shown in Table 2.
[0074] Table 2 Overall results of UniGReM comparison experiment
[0075]
[0076] From the experimental results, it can be found that the small-sample unified granularity relationship extraction model based on the large language model proposed in the present invention significantly exceeds the overall evaluation indicators of the comparison models TANL and seq2rel on four data sets of three types of relationship extraction tasks, achieving the best effect, thereby verifying the effectiveness of the method in the unified granularity relationship extraction task.
[0077] To evaluate the accuracy of the proposed small-sample unified granularity relation extraction model based on a large language model in small-sample scenarios, we randomly sampled small-sample datasets of 100, 200, 10%, 20%, and 30% of the full Re-DocRED dataset. The experimental results of our method on these small-sample datasets are shown in Table 3.
[0078] Table 3 UniGReM small sample experimental results
[0079]
[0080]
[0081] Experimental results show that on small sample datasets ranging from 100 to 900 samples, the proposed model UniGReM shows improvements in the F1 metric compared to the comparison models TANL and seq2rel. This phenomenon demonstrates that the proposed model surpasses the comparison models in accuracy across all small sample datasets, effectively improving the accuracy of the model in document-level relationship extraction tasks in small sample scenarios.
[0082] To evaluate the impact of various algorithmic modules in our proposed unified granularity relation extraction method for small sample scenarios based on a large language model, we conducted ablation experiments to verify the effects of context learning, thought chaining, and entity-enhanced context on our method. The ablation results are shown in Table 4.
[0083] Table 4 UniGReM ablation experiment results
[0084]
[0085] Among them, "-w / o taskDes" is a variant model that removes the task description information of the prompt template; "-w / oICL" represents a variant model after removing the context learning demonstration example in the prompt template; "-w / oentityEnhance" represents a variant model that does not use the entity enhanced context but only uses the original context; "-w / o CoT" represents a variant model that removes the thought chain information in the output text and only outputs serialized relationship triples; "-w / o ALL" means that the input template only contains entities and unenhanced context, and the output does not contain CoT information, but only contains relationship triples. From the experimental results, it can be found that each algorithm module can improve the relationship extraction accuracy of the model of the present invention.
Claims
1. A small sample unified granularity relationship extraction method based on a large language model, comprising the following steps: (1) Task information description: provide a specific task description as part of the input context of the large language model; (2) Contextual example demonstration: Give the large language model a comparable example as a contextual demonstration and incorporate it into the input in the form of linear natural language; (3) Contextual entity enhancement: Entity enhancement is performed on the context input to the large language model, thereby better indicating the location information of the large language model entity in the context; (4) Thinking chain output defines a way to serialize relation triples and integrates the thinking chain reasoning information into it, thereby transforming the relation extraction problem into a text generation problem.
2. The method for extracting small sample uniform granularity relationships based on a large language model according to claim 1, characterized in that: In the step (1), a unified relationship extraction task definition is given that can encompass three relationship extraction tasks at the sentence level, document level and cross-document level.
3. The method for extracting small sample uniform granularity relationships based on a large language model according to claim 1, characterized in that: In the step (2), the context demonstration example is generated by a randomly selected sample from the training set of the given data set through the prompt template constructor.
4. The method for extracting small sample uniform granularity relationships based on a large language model according to claim 1, characterized in that: In the step (3), entity enhancement only adds an identifier at the end of the entity.
5. The method for extracting small sample uniform granularity relationships based on a large language model according to claim 1, characterized in that: In step (4), the infix form of the relation triple is converted to a suffix form. First, the first mention of the head entity is concatenated with the type of the entity to which it belongs, and then the sentence number in which the entity appears in the context is added. The sentence number starts from 0, which helps the large language model focus on the sentences related to the entity, thereby improving the reasoning ability of relation extraction; followed by a similar tail entity representation; finally, the relation representation surrounded by "@" is output.
Citation Information
Cited By
Self-adaptive cross-document entity relation extraction method and device based on path construction, equipment and storage medium
CN120893437A
Adaptive cross-document entity relation extraction method, device and equipment based on path construction and storage medium
CN120893437B